AI Business Strategist
D1 · AI Fundamentals and Literacy
The business vocabulary of AI, ML and GenAI, data types and quality, training on historical data, global standards, rule-based versus AI, agents, monitoring, shadow AI, prompting, tokens and model adaptation.
This domain carries 24% of the blueprint — roughly 20 of 85 items on our mock. It is not a technical exam: it tests whether a business leader has enough literacy to hold a straight conversation with a technical team, choose the right kind of solution, and know what to ask for and what to worry about. You will not be asked to derive an algorithm; you will be asked which vocabulary applies, when AI is the wrong tool, what an agent’s autonomy costs you in oversight, and why a model that worked in January quietly degrades by June. The whole domain rewards one habit: understanding AI well enough to make a decision, without pretending to be an engineer.
What you need to know
AI is the umbrella; machine learning is AI that learns patterns from data; generative AI is ML that produces new content. A model is trained on historical data, then used for inference to make predictions — so it inherits whatever the training data encoded, including yesterday’s bias. Data type (structured versus unstructured) and data quality drive outcomes more than model choice does. Rule-based automation beats AI whenever the logic is deterministic and auditable; AI earns its place where patterns are too varied to enumerate. Agents add autonomy and tool use, which multiplies both value and governance exposure. Because the world moves and models do not, every AI solution needs ongoing monitoring for drift, and every organisation needs a transparent tool-classification list to keep shadow AI in check.
Learning objectives
By the end of this page you should be able to:
- Use the core business vocabulary of AI correctly — algorithm, model, training, inference, prediction — and place AI, ML and GenAI in their nested relationship.
- Distinguish structured from unstructured data and explain why data type and data quality drive AI outcomes.
- Explain what training on historical data implies for a business: representativeness, recency, feedback loops and inherited bias.
- Describe at a business level what ISO/IEC 23053 and ISO/IEC 42001 are for.
- Decide when rule-based automation is the correct tool and when AI is the wrong one.
- Distinguish AI agents from other AI solutions and reason about the governance consequences of autonomy.
- Specify the ongoing monitoring a business owner must demand: model drift, data drift and performance decay.
- Establish a transparent AI-tool classification (approved / blocked / under evaluation) to mitigate shadow-AI risk.
- Apply basic prompt-engineering principles and explain token and context-window limits as business constraints.
- Choose between prompting, RAG and fine-tuning on cost, effort, freshness and control.
Task statements covered
| Official task / skill | Where this page teaches it |
|---|---|
| 1.1.1 Core AI concepts in business contexts | 1.1 |
| 1.1.2 Distinguish AI, ML and GenAI | 1.1 |
| 1.1.3 Structured vs unstructured data | 1.2 |
| 1.1.4 Why data quality matters | 1.2 |
| 1.1.5 Training a model on historical data | 1.3 |
| 1.1.6 Global frameworks and unified vocabulary (ISO/IEC 23053, ISO/IEC 42001) | 1.4 |
| 1.2.1 Rule-based automation vs AI | 1.5, Decision framework |
| 1.2.2 AI agents vs other AI solutions; agent capabilities | 1.6 |
| 1.2.3 Ongoing monitoring; model drift and performance changes | 1.7 |
| 1.2.4 Transparent tool classification; shadow-AI risk | 1.8 |
| 1.3.1 Basic prompt-engineering principles | 1.9 |
| 1.3.2 Token limits and context-window constraints | 1.10 |
| 1.3.3 Model adaptation (RAG, fine-tuning) | 1.11 |
1.1 AI, ML, GenAI and the words that trip leaders up
The exam expects you to use five words precisely and to place three fields in their nested order. Getting the vocabulary wrong is the fastest way to lose a technical team’s confidence — and to pick a wrong answer.
| Term | Plain-English meaning | Business analogy |
|---|---|---|
| Algorithm | The method or recipe for learning or deciding | The training regime, not the athlete |
| Model | The trained artefact that makes predictions | The athlete after training |
| Training | Fitting the model to historical data | The months of practice |
| Inference | Using the trained model on new input | Race day |
| Prediction | The model’s output for a given input | The result of one race |
The three fields are nested, not rival:
┌──────────────────────────────────────────────┐ │ ARTIFICIAL INTELLIGENCE │ │ systems that perform tasks needing │ │ human-like judgement │ │ ┌──────────────────────────────────────┐ │ │ │ MACHINE LEARNING │ │ │ │ learns patterns from data instead │ │ │ │ of being explicitly programmed │ │ │ │ ┌──────────────────────────────┐ │ │ │ │ │ GENERATIVE AI │ │ │ │ │ │ ML that generates new content │ │ │ │ │ │ (text, image, audio, code) │ │ │ │ │ └──────────────────────────────┘ │ │ │ └──────────────────────────────────────┘ │ └──────────────────────────────────────────────┘A recommendation engine is ML but not generative. A rules engine that flags transactions over a threshold is neither ML nor GenAI — it is deterministic automation. GenAI is a subset of ML, which is a subset of AI; the exam punishes the reflex to call everything “AI”.
Exam signal
Stems that say “the system learns from examples” point to ML; “generates a draft / summary / image / new content” points to GenAI; “applies a fixed set of rules” points to rule-based automation, not AI at all. Words like train, infer, predict are being used precisely — read them literally.
1.2 Structured vs unstructured data, and why quality drives outcomes
Data type shapes which solution is even possible, and data quality caps how good the outcome can be regardless of the model.
| Structured data | Unstructured data | |
|---|---|---|
| Shape | Rows and columns, fixed fields | Free text, images, audio, video, documents |
| Examples | Transactions, CRM records, sensor readings | Emails, contracts, call recordings, photos |
| Classic fit | Classical ML, BI dashboards | GenAI, computer vision, document extraction |
| Business note | Easy to query and audit | The 80% of enterprise data that was previously hard to use |
The leadership point is that GenAI made unstructured data commercially useful — the contracts, tickets and call transcripts that used to sit unread. But every kind of AI is bounded by data quality. A useful business checklist for quality is complete, correct, current, consistent, representative:
DATA QUALITY → OUTCOME CEILING ───────────────────────────────── Incomplete → gaps the model fills with guesses Incorrect → confident wrong answers Stale → right answer to last year's question Inconsistent → the same entity treated three ways Unrepresentative → works for the majority, fails a segmentWorked example. A support team wants to auto-triage tickets. The model reaches 91% accuracy on the majority language but 62% on a customer segment that writes in a second language under-represented in the training data. Headline accuracy of 88% hides a segment failure that is also the highest-complaint segment. The fix is a data decision (represent the segment), not a model swap — quality and representativeness, not cleverness, set the ceiling.
1.3 Training on historical data — and what it implies
A model trained on historical data is, by construction, a compressed opinion about the past. Three consequences follow, and each is a business risk, not a technical footnote.
- Representativeness. The model is only as fair as the data was balanced. If a hiring dataset over-represents one profile, the model learns that profile as “good”.
- Recency. The world moves; the training snapshot does not. A demand model trained before a market shift keeps answering last year’s question.
- Feedback loops. When a model’s own outputs become tomorrow’s training data, small biases compound. Recommend the same products, gather clicks on those products, recommend them harder.
The most examined line: yesterday’s data encodes yesterday’s bias. A model does not invent discrimination; it faithfully reproduces whatever the historical data contained, then applies it at scale and at speed. That is why a business owner asks what period is this trained on, who is in the data, who is missing, and how will the model’s own outputs feed back in before trusting a prediction.
HISTORICAL DATA ──train──► MODEL ──inference──► DECISIONS ▲ │ └────────── outputs become new data ─────────┘ (feedback loop: bias can compound)Exam signal
Stems about “the model was trained on data from before / from one region / from one customer type”, or “performance was fine at launch but has degraded”, are testing representativeness, recency or drift. The correct answer questions the data, not the algorithm.
1.4 Global frameworks and a unified vocabulary
The exam asks you to be aware of two ISO/IEC standards at a business level — what each is for, not their clause numbers.
| Standard | What it is for, in business terms |
|---|---|
| ISO/IEC 42001 | An AI management system standard — the certifiable, auditable “how we run AI responsibly as an organisation” standard, comparable in spirit to a quality or information-security management system. AWS has stated support for it. |
| ISO/IEC 23053 | A framework for AI systems that use machine learning — a shared reference model and vocabulary for describing how ML systems are structured. |
Why a strategist cares: a common vocabulary lets legal, risk, engineering and the business argue about the same thing, and a recognised management-system standard turns “trust us, we’re responsible” into something a customer or regulator can audit. You will also meet the NIST AI Risk Management Framework (Govern, Map, Measure, Manage) and risk-tiered regulation such as the EU AI Act in Domain 3 — the fundamentals here are that standards exist, they give everyone the same words, and one of them (42001) is a certifiable management system.
1.5 Rule-based automation versus AI
The single most valuable literacy skill on this exam is knowing when not to reach for AI. Rule-based automation is deterministic, transparent, cheap to audit and correct every time the rule is correct. AI is probabilistic, powerful on messy patterns, and wrong some of the time by design.
| Dimension | Rule-based automation | AI / ML |
|---|---|---|
| Logic | Explicit, human-written rules | Patterns learned from data |
| Determinism | Same input → same output, always | Probabilistic; can vary |
| Auditability | Fully explainable by reading the rule | Needs explainability tooling |
| Best when | Logic is stable, known, regulated | Patterns are too varied to enumerate |
| Fails when | Rules explode combinatorially | Data is poor, sparse or unrepresentative |
| Cost of error | Predictable | Sometimes silent and confident |
Honest examples where AI is the wrong tool:
- Calculating tax or interest — a formula, deterministic, must be exact and auditable. Use a rule.
- Enforcing a compliance threshold (“block any transfer over the limit”) — a rule, not a model.
- Routing a form by a checkbox value — a rule.
- Any decision that must be identical and defensible every time — a rule, because “usually right” is not good enough.
AI earns its place where the input is unstructured or the patterns are too varied to write down: summarising a call, extracting fields from a messy invoice, flagging anomalies no one pre-listed, drafting a first version. When a stem describes a deterministic, well-specified, high-stakes-must-be-exact problem, the sophisticated answer is a rule.
1.6 AI agents versus other AI solutions
An agent is an AI solution that can take actions toward a goal, not just produce an answer. The defining capabilities are autonomy (it decides the next step), tool use (it calls systems, APIs or functions), agent-to-agent communication (agents coordinate) and orchestration (a controller sequences the steps). A plain GenAI assistant answers; an agent does.
ASSISTANT AGENT ───────── ───── You → prompt → answer Goal → plan → act (tool) → observe You act on the answer → decide next step → repeat → until goal met or stopped Human in every loop Human sets the goal and the guardrailsThe leadership point is that autonomy and governance move together. The more an agent decides and acts on its own — spending money, sending messages, changing records — the more it can accomplish and the more damage a single wrong step can do. So autonomy is a governance dial, not a free upgrade:
| Autonomy level | Value | Governance consequence |
|---|---|---|
| Suggests, human executes | Modest | Low; human is the gate |
| Acts in a sandbox, human approves | Higher | Approval gate, reversible |
| Acts on real systems within limits | High | Needs hard guardrails, logging, escalation |
| Fully autonomous, chains actions | Highest | Highest blast radius; needs kill-switch and audit |
Exam signal
Stems with “autonomously”, “takes actions”, “calls tools / systems”, “chains steps”, “coordinates with other agents” are agent items. When they add “spends / sends / changes / deletes”, the correct answer pairs the capability with oversight: approval gates, limits, logging and an escalation path — never “let it run because it is faster”.
1.7 Why AI needs ongoing monitoring
A trained model is a snapshot; the world keeps moving. That gap is why the correct answer to “we deployed it, are we done?” is always “no”. Three related decays matter, and a business owner must be able to name what they are asking for.
| Phenomenon | What changes | Business symptom |
|---|---|---|
| Data drift | The input distribution shifts (new customers, new phrasing, new products) | Model sees inputs unlike its training data |
| Concept / model drift | The relationship between input and correct answer shifts | Yesterday’s right answer is today’s wrong one |
| Performance decay | Accuracy, precision or business KPI erodes over time | Metrics that were fine at launch slide |
What a business owner must ask for, in plain terms: a baseline captured at launch; ongoing monitoring against it; alert thresholds that page a human; a defined cadence for review and retraining; and an owner accountable for acting on the alert. On AWS, tooling such as SageMaker Model Monitor exists to detect and alert on inaccurate predictions from deployed models — but the exam only needs you to know that monitoring is not optional and someone must own it. “Set it and forget it” is always the wrong answer.
1.8 Shadow AI and transparent tool classification
Shadow AI is employees using AI tools that the organisation has not vetted — pasting confidential data into a personal chatbot to get work done faster. It is rarely malicious; it is usually a symptom of a governed path that is too slow or does not exist. The mitigation the exam wants is transparent classification: a living, published list of tools in three states.
| State | Meaning | Example criterion |
|---|---|---|
| Approved | Vetted for stated data classes and use cases | Passed data, security and terms review |
| Blocked | Not permitted; a safe alternative is named | Fails no-training terms or residency needs |
| Under evaluation | Being assessed; do not use for real data yet | Trial in progress; decision date set |
Two things make classification work rather than perform caution: it is transparent (people can see the list and why a tool is where it is) and it comes with a fast approved default, so the governed path is easier than the ungoverned one. A blanket ban with no approved alternative does not stop shadow AI; it guarantees it. This is the fundamentals-level version of a theme Domain 3 develops into a full governance operating model.
1.9 Prompt-engineering principles at a business level
You will not be tested on clever prompt tricks, but on the principle that a model’s output quality is largely a function of the instruction it was given. The business-level checklist:
- Role — tell the model who to be (“you are a claims analyst”).
- Task — state the single job precisely.
- Context — supply the facts the model needs instead of hoping it guesses.
- Constraints — length, tone, what to avoid, what must be included.
- Format — the exact shape of the output (table, bullets, one page).
- Examples — one or two worked examples steer format and quality more than a paragraph of description.
- Success criterion — how you will judge the answer.
The leadership takeaway: vague in, vague out. When a team complains that “the AI gives generic answers”, the first question is what it was asked, not whether to buy a bigger model. Prompt clarity is the cheapest quality lever there is.
1.10 Tokens and context windows as business constraints
A token is a chunk of text (roughly a short word or word-piece) that the model reads and writes; you are typically billed per token and the model can only “see” a bounded amount of text at once — its context window.
CONTEXT WINDOW = how much the model can hold in view at once ┌───────────────────────────────────────────────────────┐ │ system + instructions │ your documents │ conversation │ ← output └───────────────────────────────────────────────────────┘ if the total exceeds the window, the oldest text falls outTwo business constraints follow. First, cost scales with tokens: a workflow that stuffs a whole 200-page manual into every request is expensive and often slower, and RAG (1.11) exists partly to avoid it. Second, the window is finite: past a limit the model forgets earlier context, truncates the document, or loses the thread of a long conversation — which shows up as a quality problem that is really a context-limit problem. When a stem says the model “forgets earlier instructions” or “misses details in a very long document”, the discriminator is the context window, not the model’s intelligence.
Exam signal
“Long document”, “forgets what we said earlier”, “cost is higher than expected per request”, “truncated” point to token and context-window constraints. The fix is usually to send less but more relevant context (RAG, summarisation, chunking) — not to buy a smarter model.
1.11 Model adaptation: prompt → RAG → fine-tuning
There is a ladder of ways to make a general model fit a specific business need. Climb it only as far as you must, because each rung costs more effort and control.
| Technique | What it does | Cost / effort | Data freshness | Control | Use when |
|---|---|---|---|---|---|
| Prompt engineering | Better instructions and context in the request | Lowest | As fresh as what you paste | Low | The base model can do it with guidance |
| RAG | Retrieves relevant company data at query time and feeds it in | Medium | Always current — update the source, not the model | Medium | Answers must reflect your changing, proprietary knowledge |
| Fine-tuning | Retrains the model’s weights on your examples | Highest | Frozen at training time; re-tune to refresh | High | You need a consistent style/behaviour that instructions cannot achieve, and the data is stable |
Retrieval-Augmented Generation (RAG) is the workhorse for business knowledge: it keeps the facts in a source you control, so updating a policy document updates the answers with no retraining. On AWS this is what managed Knowledge Bases provide at a strategic level — RAG without the infrastructure. Fine-tuning changes the model itself; it is powerful for consistent tone or specialised behaviour but expensive, slow to refresh, and the wrong tool when the underlying facts change weekly.
Need better answers from a general model? │ ├─ Better instructions / context enough? → PROMPT ENGINEERING (start here) │ ├─ Answers must reflect changing, proprietary facts? → RAG │ └─ Need consistent specialised behaviour, stable data, and prompting/RAG are not enough? → FINE-TUNING (last, most costly)The exam’s preferred instinct: start at the cheapest rung and climb only when forced. Reaching straight for fine-tuning when a prompt or RAG would do is the classic over-engineering distractor.
Decision framework
The capability-to-solution selector
When a business problem lands on your desk, this selector routes it to the kind of solution before anyone talks about a specific product. Score the problem on four axes, then read the recommendation.
| Axis | Question | Leans deterministic ↔ leans AI |
|---|---|---|
| Determinism | Is the correct answer fixed and specifiable as rules? | Yes → rules · No → AI |
| Data availability | Do we have quality, representative data on the pattern? | No → rules/defer · Yes → ML/GenAI |
| Error tolerance | Can the process absorb an occasional wrong answer? | No → rules or heavy oversight · Yes → AI |
| Auditability need | Must every decision be explained and identical? | Yes → rules (or explainable ML + review) · No → GenAI/agentic OK |
How to read it:
| Profile | Recommended solution class |
|---|---|
| Deterministic, must be exact, fully auditable | Rule-based automation |
| Patterns in structured data, prediction/classification, some error tolerance | Classical ML |
| Unstructured input, generate/summarise/extract, human checks output | Generative AI |
| Multi-step goal, needs to act across systems, tolerable and reversible errors, guardrails in place | Agentic AI (with oversight matched to autonomy) |
Worked application. A lender wants to (a) decide loan eligibility against a fixed regulatory rule, (b) predict default risk, and (c) draft the customer decision letter.
- (a) Eligibility is deterministic, high-stakes, must be identical and auditable → rule-based automation. Using a model here would be both riskier and harder to defend.
- (b) Default risk is a pattern in structured historical data with error tolerance managed by policy → classical ML, with explainability (e.g. SageMaker Clarify at the strategic level) and human review of edge cases because the domain is regulated.
- (c) The letter is unstructured generation from a template and the facts → GenAI, with a human review gate before anything goes to the customer.
One business problem, three different solution classes. The framework’s value is refusing the reflex to make all three “an AI project”.
Common mistakes
| Mistake | Why it happens | What to do instead |
|---|---|---|
| Calling every automation “AI” | The word is fashionable and vague | Reserve AI/ML for learned-pattern systems; call deterministic logic what it is |
| Reaching for AI on a deterministic problem | AI feels modern and impressive | If the answer is a fixed rule, use a rule — it is exact, cheap and auditable |
| Trusting a model because its output is fluent | Confidence reads as competence | Fluency is not accuracy; a model can be confidently wrong |
| Ignoring data quality and blaming the model | The model is the visible part | Fix representativeness, recency and correctness first; quality caps the ceiling |
| Assuming a launched model stays accurate | “We deployed it, we’re done” | Demand baselines, drift monitoring, a review cadence and an owner |
| Treating an agent’s autonomy as a free upgrade | More autonomy looks like more value | Match oversight to autonomy; more action power means more guardrails |
| Banning AI tools outright to stop shadow AI | A ban feels safe and simple | Publish approved/blocked/under-evaluation and provide a fast approved default |
| Jumping straight to fine-tuning | It sounds like the “real” solution | Start with prompting, then RAG; fine-tune only when forced |
| Stuffing whole documents into every prompt | “Give it everything to be safe” | Send relevant context (RAG, summaries); tokens cost money and windows are finite |
| Believing bias comes from the algorithm | The maths feels neutral | Bias usually enters through historical data; interrogate who is in and missing |
| Confusing RAG with fine-tuning | Both “customise the model” loosely | RAG feeds current data at query time; fine-tuning retrains weights on stable data |
| Expecting the model to guess missing context | It answers so fluently anyway | Supply role, task, context, constraints and format; vague in, vague out |
Scenario walkthrough
Scenario. A retailer’s operations director asks you to “add AI” to three problems at once. First, a returns-eligibility check: the policy is a fixed set of rules (within 30 days, unworn, receipt present). Second, demand forecasting for seasonal stock, where two years of clean sales history exist. Third, an auto-responder that will read a customer’s free-text complaint and send a resolution email — refund, replacement or apology — on its own to cut response time. A vendor has pitched a single “AI platform” for all three, and a fine-tuned model “trained on your data” for the responder. The director wants a recommendation by Friday and is impressed by the fine-tuning pitch.
Expert reasoning trace.
-
Refuse the one-size framing. These are three different problems on the capability-to-solution selector, and bundling them into one “AI project” is the first error. I score each separately.
-
Returns eligibility is deterministic and auditable — the answer is fixed by policy and must be identical every time and defensible to a customer or regulator. That is rule-based automation, not AI. Using a probabilistic model here would introduce variability into a rule that has no reason to vary, and make disputes harder to defend. The sophisticated answer says “this one is not an AI problem”.
-
Demand forecasting is a pattern in structured historical data with error tolerance managed by safety stock. That is classical ML. But I immediately flag recency and representativeness: two years of history that predate a market shift will forecast last year’s demand, so I ask what period the data covers and require drift monitoring, because a forecast that was accurate at launch decays as conditions move.
-
The auto-responder is the dangerous one. It is GenAI (unstructured input, generated output) plus autonomy (it sends, and refunds are money leaving the business). Autonomy and governance move together: an agent that spends money on its own has a large blast radius. So the recommendation is GenAI to draft, with a human approval gate before anything is sent or refunded — at least until monitored performance justifies raising autonomy on low-value, reversible cases.
-
Challenge the fine-tuning pitch. The responder’s facts — this customer, this order, current policy — change constantly, and fine-tuning freezes knowledge at training time. The right adaptation is RAG over current order and policy data, not fine-tuning, because updating a policy document must update the answers without a retrain. Fine-tuning here would be the most expensive rung of the ladder solving a problem the cheaper rung solves better.
-
Name the monitoring and ownership. Each AI component gets a baseline, drift monitoring and a named owner; the deterministic returns check does not need drift monitoring because it does not learn.
Exam-correct recommendation: rule-based automation for returns eligibility; classical ML with drift monitoring for forecasting; GenAI with RAG and a human approval gate for the responder, raising autonomy only for low-value reversible cases once monitored. Reject the single-platform bundle, reject fine-tuning for the responder, and reject any full autonomy that lets the system refund money unsupervised on day one.
Exam traps in this domain
| Trap | Why it is tempting | The discriminator |
|---|---|---|
| “It’s automation, so it’s AI” | The terms blur in marketing | AI/ML learns patterns; a fixed rule is deterministic automation |
| “Use AI because it’s more advanced” | Modern beats old-fashioned | If the logic is deterministic and must be exact, a rule is the superior answer |
| “The model is fluent, so it’s right” | Confidence signals competence | Fluency is a language property, not a truth signal |
| “Deployed and accurate at launch, so we’re done” | Launch feels like the finish line | Drift and decay are guaranteed; monitoring and an owner are required |
| “Fine-tune it on our data” for changing facts | It sounds like the serious option | Fine-tuning freezes knowledge; RAG keeps it current at query time |
| “Give it full autonomy to be fast” | Speed is attractive | Autonomy scales blast radius; oversight must scale with it |
| “Ban all AI tools to be safe” | A ban feels decisive | Bans drive shadow AI; classify tools and offer a fast approved default |
| “Bias is a technical/algorithm problem” | The maths seems neutral | Bias mostly enters via historical data; interrogate the data, not just the model |
| “Buy a bigger model to fix generic answers” | Bigger sounds better | Vague prompts and thin context cause generic output; fix the instruction first |
| “One AI platform for every problem” | Simplicity and one vendor | Different problems need different solution classes; route each on the selector |
Practice questions
Each item states how many responses to select. Attempt before revealing.
Q1 · A business analyst describes a system that 'learns patterns from historical transactions to flag likely fraud'. Which term describes it MOST precisely? (Select one)
A. Generative AI, because it produces an output. B. Machine learning, because it learns patterns from data rather than following fixed rules. C. Rule-based automation, because it flags transactions. D. An AI agent, because it takes an action.
Answer: B. Learning patterns from data is the definition of machine learning; it is not generative because it classifies rather than creates content (A), it is not a fixed-rule system (C), and flagging is not autonomous multi-step action (D).
Q2 · A returns policy says 'refund if returned within 30 days, unworn, with a receipt'. A leader wants to automate the eligibility decision. What is the MOST appropriate solution class? (Select one)
A. Generative AI to interpret each request. B. A fine-tuned model trained on past returns. C. Rule-based automation, because the logic is deterministic and must be identical and auditable. D. An autonomous agent that decides case by case.
Answer: C. The policy is a fixed, deterministic rule that must be exact and defensible, so a rule engine is superior. GenAI (A) and a fine-tuned model (B) add needless variability and cost, and an autonomous agent (D) is over-engineering for a checkbox decision.
Q3 · A GenAI support model scored 88% overall accuracy but only 62% on customers writing in a second, under-represented language. What is the ROOT cause and BEST fix? (Select one)
A. The algorithm is flawed; switch models. B. Unrepresentative training data; improve representation of that segment in the data. C. The context window is too small; shorten prompts. D. Nothing; 88% overall is acceptable.
Answer: B. A segment failure driven by under-representation is a data-quality/representativeness problem, fixed by representing the segment, not by a model swap (A) or a context change (C). Accepting the headline (D) ignores that the failing segment may be the highest-stakes one.
Q4 · A demand model was accurate at launch but has steadily lost accuracy over six months, with no code changes. What is the MOST likely explanation? (Select one)
A. The model became less intelligent. B. Drift: the input distribution or the input-to-outcome relationship has shifted since training. C. The context window shrank. D. The prompts got worse.
Answer: B. Gradual post-launch decay with no code change is the signature of data or concept drift as the world moves away from the training snapshot. Models do not degrade in intelligence (A), context windows do not shrink on their own (C), and a predictive model is not prompt-driven (D).
Q5 · Which pair correctly matches the term to its meaning? (Select two)
A. Inference — using a trained model on new input to produce a prediction. B. Training — fitting the model to historical data. C. Algorithm — the trained artefact that makes predictions. D. Prediction — the recipe used to learn. E. Model — the method for learning.
Answer: A and B. Inference is applying the trained model to new input, and training is fitting the model to historical data. C, D and E swap the definitions: the model is the trained artefact and the algorithm is the learning method, while a prediction is an output, not a recipe.
Q6 · A vendor pitches a fine-tuned model for a chatbot that must answer from your frequently updated policy documents. What is the BETTER adaptation approach and why? (Select one)
A. Fine-tuning, because it bakes the knowledge into the model. B. RAG, because it retrieves current policy at query time so updates need no retraining. C. Prompt engineering alone, ignoring the documents. D. A larger context window with no retrieval.
Answer: B. RAG keeps changing facts in a source you control and always retrieves the current version, whereas fine-tuning (A) freezes knowledge at training time and needs re-tuning to refresh. Prompting alone (C) has no access to the documents, and a bigger window (D) does not solve retrieval of the right passages.
Q7 · An organisation discovers staff pasting confidential data into personal chatbots. Which response BEST mitigates shadow AI? (Select one)
A. Ban all AI tools company-wide. B. Publish an approved/blocked/under-evaluation classification and provide a fast approved default tool. C. Ignore it; the tools are convenient. D. Monitor employees covertly.
Answer: B. Transparent classification plus a fast approved path removes the reason people go rogue. A blanket ban (A) drives shadow AI underground rather than stopping it, ignoring it (C) leaves the data risk, and covert monitoring (D) addresses trust, not the missing governed path.
Q8 · An agent will autonomously issue customer refunds based on complaint emails. Which TWO controls are MOST important before go-live? (Select two)
A. A human approval gate for refunds above a defined value. B. Logging of every action with an escalation path. C. Removing all constraints so it responds fastest. D. Choosing the newest model regardless of fit. E. Turning off monitoring to reduce cost.
Answer: A and B. Autonomy that spends money needs an approval gate for higher-value cases and full logging with escalation, because the blast radius is money leaving the business. Removing constraints (C), chasing the newest model (D) and disabling monitoring (E) all increase risk rather than control it.
Q9 · A leader complains that 'the AI keeps giving generic answers' and wants to buy a more powerful model. What is the MOST cost-effective first step? (Select one)
A. Buy the larger model immediately. B. Improve the prompt with role, task, context, constraints and format before spending on a bigger model. C. Fine-tune the model on company data. D. Reduce the context window.
Answer: B. Generic output usually reflects a vague prompt and thin context, so improving the instruction is the cheapest, fastest lever. Buying a bigger model (A) or fine-tuning (C) is premature spend, and shrinking the context window (D) would remove useful information.
Q10 · What is ISO/IEC 42001, at a business level? (Select one)
A. A pricing framework for AI services. B. An AI management-system standard — a certifiable, auditable way to run AI responsibly as an organisation. C. A specific machine-learning algorithm. D. A US federal law.
Answer: B. ISO/IEC 42001 is a management-system standard for governing AI responsibly and can be certified against, which is why it matters to a strategist. It is not a pricing model (A), an algorithm (C) or a national law (D).
Q11 · Which statement about training on historical data is MOST accurate for a business leader? (Select one)
A. Models invent their own biases independently of the data. B. A model reproduces patterns in its training data, so unrepresentative or outdated data leads to biased or stale predictions. C. Historical data guarantees accurate future predictions. D. Once trained, a model is immune to changes in the world.
Answer: B. A model faithfully reproduces the historical data, including its biases and its age, then applies them at scale. It does not invent bias from nothing (A), the past does not guarantee the future (C), and a trained model is exactly what drifts as the world changes (D).
Q12 · A team stuffs an entire 300-page manual into every request and complains costs are high and the model 'misses details'. Which explanation fits BEST? (Select one)
A. The model is not intelligent enough. B. Token and context-window constraints: cost scales with tokens and very long inputs can exceed or crowd the window; retrieve relevant sections instead. C. The prompt temperature is wrong. D. The data is unstructured.
Answer: B. Billing per token and a finite context window explain both the cost and the missed details, and RAG or chunking sends only relevant passages. Model intelligence (A), temperature (C) and data structure (D) do not explain per-request cost and context overflow.
Q13 · Which are examples where rule-based automation is the CORRECT choice over AI? (Select two)
A. Calculating interest to the exact cent for a statement. B. Blocking any transaction above a fixed regulatory threshold. C. Summarising an unstructured customer call. D. Extracting fields from messy scanned invoices. E. Drafting a marketing email.
Answer: A and B. Exact interest calculation and a fixed-threshold block are deterministic, must be exact and auditable, and are perfect for rules. Summarising calls (C), extracting from messy documents (D) and drafting copy (E) involve unstructured input or generation where AI earns its place.
Q14 · How do AI agents differ from a plain GenAI assistant? (Select one)
A. Agents are always more accurate. B. Agents can act autonomously toward a goal, using tools and orchestrating multiple steps, whereas an assistant produces an answer for a human to act on. C. Agents never need human oversight. D. Assistants cannot use any external data.
Answer: B. The defining difference is autonomous, tool-using, multi-step action toward a goal. Agents are not inherently more accurate (A), they need more oversight as autonomy rises (C), and assistants can use supplied data (D).
Q15 · A business owner is told a deployed model 'runs itself'. What should the owner insist on? (Select one)
A. Nothing; a good model needs no maintenance. B. A launch baseline, ongoing drift monitoring against it, alert thresholds, a review/retrain cadence and a named owner. C. Immediate replacement with a newer model. D. Turning off logging to save cost.
Answer: B. Every learned model drifts, so monitoring, thresholds, a cadence and an accountable owner are mandatory. ‘Runs itself’ (A) is the trap, a newer model (C) drifts too, and disabling logging (D) removes the very signal needed to detect problems.
Q16 · Which best distinguishes structured from unstructured data, and why it matters? (Select one)
A. Structured data is always more valuable. B. Structured data fits rows and columns and suits classical ML and BI; unstructured data (text, images, audio) is where GenAI unlocked previously unusable enterprise content. C. Unstructured data cannot be used by AI at all. D. The distinction has no bearing on solution choice.
Answer: B. The type of data shapes which solution is viable — GenAI made the large store of unstructured content usable. Structured is not inherently more valuable (A), unstructured is very much usable (C), and the distinction directly drives solution choice (D).
Q17 · A model's own recommendations increasingly narrow what customers see, and the narrowed behaviour becomes next month's training data. What phenomenon is this? (Select one)
A. A healthy optimisation loop. B. A feedback loop that can compound bias, because the model’s outputs shape the data it later learns from. C. A context-window overflow. D. Fine-tuning.
Answer: B. When outputs become future training data, small biases can amplify over time — a feedback loop, not a benign optimisation (A). It is unrelated to context windows (C) and is not the deliberate retraining that fine-tuning describes (D).
Q18 · A lender must (a) apply a fixed eligibility rule, (b) predict default risk from structured history, and (c) draft the decision letter. Using the capability-to-solution selector, which mapping is CORRECT? (Select one)
A. GenAI for all three, on one platform. B. Rule-based automation for (a), classical ML with explainability and review for (b), GenAI with a human review gate for (c). C. Fine-tuned model for all three. D. Autonomous agent for all three.
Answer: B. The three problems have different profiles: deterministic rule, pattern-in-structured-data prediction, and unstructured generation. A single GenAI platform (A), one fine-tuned model (C) or a blanket agent (D) all ignore that different problems need different solution classes.
Q19 · Which two questions should a business owner ask about the DATA behind a model whose predictions they will rely on? (Select two)
A. What time period was it trained on, and is that still current? B. Who is represented in the data and who is missing? C. What colour is the dashboard? D. How many GPUs were used? E. What is the vendor’s logo?
Answer: A and B. Recency and representativeness are the data questions that determine whether predictions are trustworthy and fair. Dashboard colour (C), GPU count (D) and vendor branding (E) tell you nothing about outcome quality.
Q20 · A colleague argues that lowering a model's randomness setting will 'eliminate hallucinations'. What is the accurate business-level view? (Select one)
A. Correct; low randomness guarantees truth. B. It reduces variability but does not guarantee grounding; ungrounded output can still be confidently wrong, so verification and RAG matter more. C. It increases hallucinations. D. Randomness has no effect on output at all.
Answer: B. Lower randomness makes output more consistent, not more true; grounding and verification address hallucination. It does not guarantee truth (A), does not increase hallucination as a rule (C), and randomness clearly does affect output (D).
Q21 · When is fine-tuning the MOST appropriate adaptation technique? (Select one)
A. When the underlying facts change every week. B. When you need consistent specialised behaviour or style, the data is relatively stable, and prompting and RAG are not enough. C. Whenever any customisation is wanted, as the first step. D. To reduce token costs on a single query.
Answer: B. Fine-tuning suits stable, specialised behaviour that instructions and retrieval cannot achieve. It is wrong for fast-changing facts (A), it is the last rung not the first (C), and it does not exist to trim a single query’s tokens (D).
Q22 · An operations director wants 'one AI platform' to handle a deterministic returns check, a structured-data forecast and an autonomous refund responder. What is the BEST strategic response? (Select two)
A. Route each problem separately: a rule for returns, classical ML for forecasting, GenAI with RAG and a human gate for the responder. B. Match oversight to autonomy: the refund responder needs an approval gate before money leaves the business. C. Accept the single-platform bundle to simplify procurement. D. Fine-tune one model to do all three. E. Give the responder full autonomy on day one to maximise speed.
Answer: A and B. The three problems are different solution classes, and the money-spending responder must have oversight scaled to its autonomy. A single bundle (C) and one fine-tuned model (D) ignore that difference, and full day-one autonomy over refunds (E) creates an unacceptable blast radius.
Key takeaways
- AI ⊃ ML ⊃ GenAI; use algorithm, model, training, inference, prediction precisely, and do not call deterministic automation “AI”.
- Data type and data quality set the outcome ceiling; representativeness and recency matter more than model cleverness.
- Training on historical data means the model reproduces yesterday’s patterns and biases at scale — interrogate who is in the data and how fresh it is.
- Rule-based automation beats AI whenever the logic is deterministic, exact and must be auditable; knowing when not to use AI is a core exam skill.
- Agents add autonomy and tool use; oversight must scale with autonomy because action power is blast radius.
- Every learned model drifts — demand baselines, monitoring, thresholds, a review cadence and an owner; “set and forget” is always wrong.
- Beat shadow AI with a transparent approved/blocked/under-evaluation list plus a fast approved default, not a blanket ban.
- Climb the adaptation ladder only as far as forced: prompt first, then RAG for changing facts, and fine-tune last for stable, specialised behaviour.
Last updated Sep 18, 2026