# Six Board-Level Case Studies

Six long-form organisational scenarios worked end to end – rule-based versus AI, prioritisation and ROI, responsible-AI trade-offs, shadow AI, scaling a stalled pilot, and stalled adoption – with the reasoning the hardest exam items reward.

import { Accordions, AccordionItem } from '@prosefly/astro-components';

The longest items on AIB-C01 are not definitions. They hand you a paragraph of organisational mess – a sponsor with a fixed belief, numbers that do not quite add up, a deadline – and ask what a business strategist would actually do. These six cases train that reasoning. Each is worked end to end: the situation, the decision you are asked to make, the evidence, an analysis that names the framework and shows the arithmetic, a sequenced recommendation, the weaker answers dismantled, a map to the official [task statements](/aws/aib-c01/), and three exam-style questions.

Every figure below is **illustrative** – invented for the case, internally consistent, and arithmetically checkable. No number here is an AWS price unless it appears in the [frameworks appendix](/appendix/aws/frameworks/); where a real AWS pricing structure matters, the case names the *structure* (consumption-based, seat-based, commitment-based), not a rate. These pages complement the four domain pages and the two [mock exams](/aws/aib-c01/practice-exam/); work them after your first diagnostic.

:::caution[Independent preparation]
These scenarios are ours, not AWS's. They contain no official exam questions and no guarantee of a result. AIB-C01 is a beta exam; read the current [exam guide](https://docs.aws.amazon.com/aws-certification/latest/ai-business-strategist-01/ai-business-strategist-01.html) before you book.
:::

---

## Case 1 · The contact-centre assistant that is really a knowledge-base problem

**Situation.** Meridian Retail Bank is a mid-market bank with roughly 5,000 staff and about 900 contact-centre agents across three sites. Handle time has crept up over two years and the customer-satisfaction score has slipped from 78 to 71. The Chief Customer Officer returned from a conference convinced that a generative-AI assistant, sitting beside every agent and answering questions in natural language, is the fix. She has a vendor demo booked and wants a go/no-go from you, the transformation lead, in two weeks.

The bank has tried adjacent things. A scripted chatbot on the website deflects simple balance and branch-hours queries but is switched off for anything sensitive. An internal search tool over the policy library exists but agents describe it as "faster to phone a colleague". Nobody has measured why calls take as long as they do.

You spend three days with the floor. Agents are not short of intelligence or willingness; they are short of *answers they can trust*. The product and compliance policies live in four systems: a wiki last curated eighteen months ago, a shared drive of PDFs, an email folder of "clarifications" from the compliance team, and tribal knowledge. When an agent gets a hard question they either guess, place the customer on a long hold while they hunt, or escalate. A sampled audit finds that 31% of the policy answers in the wiki are stale or contradicted by a later email clarification.

The sponsor believes the assistant will "just know" the answers. She has not asked where the assistant would get them.

**What you are asked.** Recommend whether to fund the generative-AI contact-centre assistant now, and if not now, what to do instead and in what order.

**The evidence you have.**

| Fact | Value (illustrative) |
| --- | --- |
| Contact-centre agents | 900 |
| Average handle time (AHT) | 9.4 minutes, up from 8.1 two years ago |
| Calls per agent per day | 46 |
| Share of AHT spent hunting for policy answers (observed) | ~22% |
| Wiki answers found stale or contradicted | 31% |
| Policy sources not consolidated | 4 separate systems |
| Regulated call types (lending, complaints, vulnerable customers) | ~18% of volume |
| Baseline CSAT | 71 (was 78) |
| Sponsor's requested decision date | 2 weeks |

**Analysis.** The right first move is to separate the *symptom* the sponsor sees from the *problem* the floor has. She sees slow, inconsistent service and reaches for a tool. The floor has an unreliable, fragmented knowledge base. A generative assistant does not create knowledge; it retrieves and rephrases what it can reach. If it is grounded in a corpus that is 31% wrong, it will answer wrong 31% of the time, fluently and at scale – and in the 18% of calls that are regulated, a confident wrong answer is a compliance event, not just a bad experience.

Apply the **solution-type ladder** the exam leans on (task 1.2.1 and 1.3.3): *rule-based automation → AI → retrieval-augmented AI (RAG)*. Deterministic, high-volume, unambiguous queries (balance, hours, reset steps) belong to rule-based automation – cheap, auditable, no drift. The remaining queries need the agent to find the *current* policy: that is a **retrieval** problem, and RAG is the pattern that grounds a model in an authoritative, current source rather than in its training data. But RAG's answer is only as good as the corpus it indexes. Governance-by-design (task 3.1.3) says fix the corpus and decide the human-oversight rule *before* you point a model at it.

The arithmetic tells the sponsor why sequencing matters. Time lost hunting per agent per day: 46 calls × 9.4 min × 22% ≈ **95 minutes**. Across 900 agents that is roughly **1,425 hours a day**. Even halving hunting time is worth ~712 hours a day – a large prize. But the prize is only real if the source of truth is trustworthy; grounding a model in a 31%-wrong corpus converts a hunting problem into a *confidently-wrong-at-scale* problem, which is worse in the regulated 18%.

Establish a **baseline before implementation** (task 2.2.2): AHT by call type, hunting time, first-contact resolution, and answer-accuracy sampled against a corrected policy set. Without it, no one can later prove the assistant helped, and "it feels faster" will not survive a board review.

**Recommendation, sequenced.**

1. **Reframe the decision** for the sponsor: the assistant is phase three, not phase one. Get agreement that the goal is trustworthy answers, not a specific tool.
2. **Consolidate and correct the knowledge base.** One authoritative source; retire the four fragmented ones; assign an owner in compliance for currency. This is the highest-leverage, lowest-risk step and it helps even the humans.
3. **Automate the deterministic tail** with rule-based flows for the unambiguous, high-volume queries.
4. **Baseline everything** (AHT by type, hunting time, FCR, accuracy) before any model touches a call.
5. **Then** pilot a RAG-grounded assistant on *non-regulated* call types, with mandatory human oversight on the regulated 18% and an escalation rule for low-confidence answers.
6. **Measure against the baseline** and decide scale/pause/terminate on evidence.

**What a weaker answer looks like and why.**

- *"Approve the assistant now; the demo was impressive."* Fluency in a demo is not grounding in your corpus. This funds a tool for a problem you have not defined and skips the baseline, so you could never prove value or catch the regulated-answer risk.
- *"Buy the assistant but only for regulated calls, since those are hardest."* Exactly backwards: the regulated calls carry the most oversight burden and the highest cost of a wrong answer. Start where a mistake is cheap.
- *"Skip AI entirely; just tell agents to search harder."* Ignores that the corpus itself is wrong; effort cannot compensate for a stale source of truth.
- *"Fine-tune a model on the policies."* Fine-tuning bakes today's policies into weights; policies change, and the model would drift out of date and be expensive to refresh. Retrieval against a maintained corpus is the fit.

**Which domains and tasks this exercises.** D1 1.2.1 (rule-based vs AI), 1.3.3 (RAG vs fine-tuning); D2 2.1.4 (when AI is *not* the answer yet), 2.2.2 (baselines); D3 3.1.3 (governance by design), 3.1.4 (human oversight for regulated calls).

**Three exam-style questions on this case.**

<Accordions>
  <AccordionItem title="Q1 · The sponsor wants a generative assistant to cut handle time, but 31% of the policy corpus is stale or contradicted. What is the BEST first action? (Select one)">
    A. Approve the assistant so agents stop hunting for answers.
    B. Consolidate and correct the knowledge base and assign an owner before grounding any model on it.
    C. Fine-tune a model on the four existing policy systems.
    D. Deploy the assistant on regulated calls first because they are hardest.

    **Answer: B.** A retrieval-grounded assistant inherits the accuracy of its source; grounding on a 31%-wrong corpus produces confident wrong answers at scale. Fixing the corpus is the highest-leverage, lowest-risk move and helps humans immediately. A funds a tool for an undefined problem. C bakes today's stale policies into weights and drifts. D starts where a mistake is most costly.
  </AccordionItem>

  <AccordionItem title="Q2 · Which model-adaptation approach best fits answering agent questions from current bank policy that changes often? (Select one)">
    A. Fine-tuning on the policy documents.
    B. Retrieval-augmented generation grounded in a maintained authoritative corpus.
    C. A larger context window with the whole wiki pasted in each time.
    D. A rule-based decision tree covering every policy.

    **Answer: B.** RAG grounds responses in a current, authoritative source and updates as the source updates – the right fit for frequently changing policy. A bakes policy into weights and goes stale. C is costly, hits context limits and still uses a stale wiki. D cannot scale to the full, evolving policy space and is unmaintainable.
  </AccordionItem>

  <AccordionItem title="Q3 · Before piloting the assistant, which TWO measures should be captured as baselines? (Select two)">
    A. Average handle time and hunting time by call type.
    B. Answer accuracy sampled against a corrected policy set.
    C. The vendor's published benchmark scores.
    D. The number of agents who attended the demo.
    E. Total contact-centre payroll for the year.

    **Answer: A and B.** Baselines must be the metrics the initiative claims to move, captured before implementation so improvement can be proven and regulated-answer risk detected. C is the vendor's number, not yours. D and E are context, not outcome baselines for this initiative.
  </AccordionItem>
</Accordions>

---

## Case 2 · Three initiatives, one budget, a board meeting next week

**Situation.** Kestrel Industrial is a manufacturer of ~12,000 staff with tight margins and a cautious board. Three AI initiatives have surfaced in the same quarter, each with an internal champion, and there is budget for one this year. The CFO wants a single recommendation at next week's board meeting, defensible on numbers, not enthusiasm.

Initiative A is **predictive maintenance** on the main production lines: forecast failures, schedule downtime, avoid unplanned stoppages. Initiative B is a **sales-copilot** that drafts proposals and answers product questions for the field sales team. Initiative C is a **document-extraction** tool that pulls fields from supplier invoices and contracts to cut manual data entry.

Each champion has a business case of varying quality. The maintenance team has real downtime cost data. The sales champion has an enthusiastic pilot but soft numbers. The finance champion has a clean, small, certain saving. The board has read that competitors are "investing heavily in AI" and does not want to fall behind, but is allergic to writing off a failed programme.

The sponsor – the COO – privately favours the sales copilot because it is the most visible. Your job is to make the recommendation the evidence supports, and to frame it so the board can act.

**What you are asked.** Recommend which single initiative to fund this year, and describe how you would treat the other two.

**The evidence you have.**

| Initiative | Annual benefit (illustrative) | One-off + annual cost | Feasibility | Data readiness | Strategic fit |
| --- | --- | --- | --- | --- | --- |
| A · Predictive maintenance | $4.0M avoided downtime | $0.6M build + $0.4M/yr run | Medium (sensor data exists) | High | High (core operations) |
| B · Sales copilot | $1.2M claimed uplift (soft) | $0.3M + $0.5M/yr | High (SaaS, quick) | Medium | Medium |
| C · Invoice extraction | $0.5M labour saved | $0.15M + $0.1M/yr | High | High | Low-medium |

**Analysis.** Score each initiative on the four criteria the exam names for prioritisation (task 2.1.3): **business value, feasibility, sustainability, strategic alignment**. Enthusiasm and visibility are not on the list; the COO's preference is a bias to manage, not a criterion.

Do the ROI arithmetic first, because the board is numbers-led. First-year ROI = (benefit − first-year cost) ÷ first-year cost, where first-year cost = one-off + one year of run.

- **A:** first-year cost $0.6M + $0.4M = $1.0M; net = $4.0M − $1.0M = $3.0M; ROI ≈ **300%**. Ongoing ROI is far higher once the build is amortised: ($4.0M − $0.4M) ÷ $0.4M = **900%**.
- **B:** first-year cost $0.3M + $0.5M = $0.8M; net = $1.2M − $0.8M = $0.4M; ROI ≈ **50%** – but the benefit is *soft*, so a realistic downside (say half the claim, $0.6M) turns it negative: $0.6M − $0.8M = **−$0.2M**.
- **C:** first-year cost $0.15M + $0.1M = $0.25M; net = $0.5M − $0.25M = $0.25M; ROI ≈ **100%**; small, certain, low strategic weight.

A is the clear leader on value *and* strategic alignment (it defends core operations, the source of the firm's margin) with high data readiness. Its only soft spot is feasibility, which is medium – so the risk to manage is delivery, not whether the prize is real.

On **investment level versus industry maturity** (task 2.3.4): in a mature, margin-thin manufacturing sector, the defensible move is to invest where AI protects the core, not to chase the most visible use case. "Competitors are investing heavily" is a reason to invest *deliberately*, not a reason to pick the shiniest option.

The other two are not simply rejected. The exam's language is **scale, pause or terminate** (task 2.1.3). C is cheap, certain and low-risk: fund it from operating budget as a quick win rather than the strategic slot, or **pause** it to next cycle. B rests on soft numbers: **pause** it and send the champion away to establish a proper baseline and a measured pilot before it competes for money again – do not terminate a plausible idea for lack of measurement.

**Recommendation, sequenced.**

1. **Fund A (predictive maintenance)** as the strategic initiative: highest value, strongest strategic fit, high data readiness, ROI ~300% year one.
2. **Manage A's feasibility risk**: stage the build, gate the spend on a proof against real line data, and set a scale/pause/terminate review at the pilot boundary.
3. **Approve C as a low-cost quick win** from operating budget if cash allows – it is nearly free and self-funding – otherwise queue it.
4. **Pause B**: return it to the champion with a requirement to establish a baseline and a measured pilot; re-evaluate next cycle on hard numbers.
5. **Present all three to the board on one page** with the four-criteria scores and the ROI maths, so the decision is legible and the COO's preference is addressed transparently.

**What a weaker answer looks like and why.**

- *"Fund the sales copilot; it is what the COO wants and it is visible."* Substitutes sponsor preference and visibility for the four criteria, and rests the case on soft numbers that go negative on a realistic downside.
- *"Fund all three at reduced scope."* Spreads a single budget across three half-resourced efforts, none of which reaches the scale needed to prove value; violates the start-with-focus logic.
- *"Terminate B and C; only A matters."* Over-prunes: C is nearly free and certain, and B is a plausible idea that only lacks measurement. Pause, do not kill.
- *"Pick C because it is the safest and cheapest."* Optimises for certainty over value and strategic alignment; the smallest, least strategic prize should not consume the strategic slot.

**Which domains and tasks this exercises.** D2 2.1.2 (build-buy-partner is implicit in each), 2.1.3 (prioritise; scale/pause/terminate), 2.2.3 (ROI), 2.3.4 (investment level vs industry maturity); D4 4.4.2 (start with short-term wins) for the treatment of C.

**Three exam-style questions on this case.**

<Accordions>
  <AccordionItem title="Q1 · With one budget and three initiatives, on what basis should the strategic slot be awarded? (Select one)">
    A. The initiative the executive sponsor personally prefers.
    B. Business value, feasibility, sustainability and strategic alignment, scored transparently.
    C. The initiative that is most visible to customers.
    D. The initiative with the lowest cost.

    **Answer: B.** These are the four prioritisation criteria; enthusiasm, visibility and raw cheapness are not among them. A substitutes sponsor bias for evidence. C picks visibility over value. D optimises cost over value and strategic fit, awarding the slot to the smallest prize.
  </AccordionItem>

  <AccordionItem title="Q2 · The sales copilot's benefit is soft and halves on a realistic downside. What is the appropriate disposition? (Select one)">
    A. Terminate it; soft numbers mean it is worthless.
    B. Fund it anyway because it is the most visible.
    C. Pause it and require a baseline and a measured pilot before it competes for budget again.
    D. Scale it immediately to gather more data in production.

    **Answer: C.** Scale/pause/terminate: a plausible idea that only lacks measurement should be paused and sent back for a baseline, not killed. A over-prunes a plausible idea. B funds on visibility. D scales an unproven case into production, the classic pilot-to-prod mistake.
  </AccordionItem>

  <AccordionItem title="Q3 · Predictive maintenance shows year-one ROI near 300% but only medium feasibility. Which TWO actions best manage this? (Select two)">
    A. Stage the build and gate spend on a proof against real line data.
    B. Set a scale/pause/terminate review at the pilot boundary.
    C. Skip the pilot to capture the benefit sooner.
    D. Reassign the budget to the safest initiative instead.
    E. Increase the benefit estimate to strengthen the board case.

    **Answer: A and B.** The prize is real and strategic; the risk is delivery, so stage the spend behind a real-data proof and a decision gate. C removes the very check that manages feasibility risk. D abandons the best prize for certainty. E inflates the case dishonestly.
  </AccordionItem>
</Accordions>

---

## Case 3 · The accurate triage model the clinical council will not approve

**Situation.** Caldera Health runs urgent-care clinics and a nurse-led telephone triage line. A data-science partner has built a triage model that recommends a disposition – self-care, book a GP, or go to emergency – from the patient's described symptoms. On a held-out test set it is **91% accurate**, better than the current protocol's measured 84%. The programme sponsor, the Chief Operating Officer, sees a way to cut triage time and handle rising demand, and wants to deploy.

The clinical governance council refuses to approve it. Their objection is not the accuracy number. It is that the model is a black box: when it recommends "self-care" for a patient a nurse would have escalated, no clinician can see *why*, and the 9% of errors are not evenly distributed – a sampled review suggests the model under-triages certain presentations more than others. In a safety-critical, regulated setting, a wrong "self-care" recommendation is the error that harms.

The sponsor argues that 91% beats 84% and that the council is being risk-averse. The data-science partner offers a more accurate but even less interpretable model. You are the AI strategy lead asked to break the deadlock.

**What you are asked.** Recommend how to proceed so that the accuracy gain is not lost and the clinical council's concerns are genuinely met.

**The evidence you have.**

| Fact | Value (illustrative) |
| --- | --- |
| Model accuracy (held-out) | 91% |
| Current protocol accuracy | 84% |
| Error distribution | Uneven; under-triage concentrated in some presentations |
| Explainability | Low – no per-decision rationale |
| Setting | Safety-critical, regulated (clinical) |
| Highest-harm error | False "self-care" for a patient needing escalation |
| Alternative offered | Higher accuracy, lower interpretability |

**Analysis.** The council is right, and the exam expects you to see why. In a safety-critical domain the relevant question is not "is aggregate accuracy higher?" but "what is the cost distribution of the errors, and can a human oversee the decision?" (tasks 3.1.1, 3.1.2, 3.1.4). Aggregate accuracy hides the error that matters: an under-triage in the highest-harm class. A model that is 91% accurate overall but concentrates its errors in "sent home a patient who needed care" is not safer than an 84% protocol whose errors are more benign.

Name the trade-off explicitly (task 3.1.2): **explainability versus raw performance**. The more accurate, less interpretable alternative moves the wrong way. In a regulated clinical setting, a clinician must be able to understand and, where needed, override a recommendation; explainability is a *requirement*, not a nice-to-have, because it is what makes human oversight real rather than nominal.

Apply a **risk-classification lens** (task 3.2.4): clinical triage that can send a patient home is high-risk by any tier. High-risk use cases carry obligations – human oversight, explainability, monitoring, documented sign-off – regardless of how good the accuracy looks. **Governance by design** (task 3.1.3) means these were requirements to build in from the start, not gates to argue about at the end.

The reframing that unlocks the deadlock: the model does not have to *replace* the nurse's judgement to add value. Deployed as **decision support with human oversight** – the model suggests, the nurse decides, with an escalation rule that any "self-care" suggestion overriding a nurse's instinct is reviewed – you capture the accuracy gain on the benign majority while a clinician backstops the harmful tail. Add the explainability and monitoring the council needs, and their objection is met on its merits.

**Recommendation, sequenced.**

1. **Reject the black-box replacement**; reject the more-accurate-less-interpretable alternative for this use case.
2. **Reclassify the use case as high-risk** and attach the obligations: human oversight, explainability, bias/error monitoring by presentation type, documented clinical sign-off.
3. **Deploy as decision support, not decision-maker**: the model recommends, a nurse decides, with a hard escalation rule on any error-prone or high-harm class.
4. **Add per-decision explanations** (or a model that provides them) so oversight is real; measure the *distribution* of errors, not just the aggregate.
5. **Monitor in production** for error drift, especially in the under-triage classes, and give the council a standing review.
6. **Bring the sponsor along** by showing that this path keeps most of the accuracy gain while making the risk defensible – deployment the council approves beats a superior model it blocks.

**What a weaker answer looks like and why.**

- *"Deploy it; 91% beats 84%."* Compares aggregates and ignores that the errors concentrate in the highest-harm class. Higher average accuracy is not higher safety.
- *"Adopt the more accurate, less interpretable model."* Optimises the wrong axis; less interpretability makes human oversight weaker in exactly the setting that most needs it.
- *"Keep the manual protocol; AI is too risky for clinical work."* Discards a real gain the majority of patients would benefit from; the answer is oversight, not abstention.
- *"Let the council approve it with a disclaimer."* A disclaimer is not oversight; it transfers blame without adding a safeguard or explainability.

**Which domains and tasks this exercises.** D3 3.1.1 (responsible-AI dimensions: explainability, safety), 3.1.2 (business goal vs responsible-AI trade-off), 3.1.3 (governance by design), 3.1.4 (human oversight and escalation), 3.2.4 (risk classification); D3 3.3.2 (uneven error/bias across a lifecycle) and 3.3.4 (reliability).

**Three exam-style questions on this case.**

<Accordions>
  <AccordionItem title="Q1 · A triage model is 91% accurate versus 84% for the manual protocol, but its errors concentrate in high-harm under-triage. What is the BEST reading? (Select one)">
    A. Deploy it; higher aggregate accuracy means higher safety.
    B. The error distribution matters more than the aggregate in a safety-critical setting; oversight and explainability are required before use.
    C. Reject AI for clinical triage entirely.
    D. Adopt the more accurate but less interpretable alternative.

    **Answer: B.** In safety-critical use the cost distribution of errors, human oversight and explainability govern the decision, not the headline accuracy. A compares aggregates and ignores the harmful tail. C throws away a real benefit for most patients. D worsens the interpretability the setting most needs.
  </AccordionItem>

  <AccordionItem title="Q2 · The sponsor offers a more accurate but less interpretable model to satisfy the council. Why is this the wrong move? (Select one)">
    A. More accurate models are always better.
    B. In a regulated clinical setting explainability enables genuine human oversight, so trading it away weakens the safeguard the council needs.
    C. The council only cares about cost.
    D. Interpretability has no bearing on oversight.

    **Answer: B.** Explainability is what makes clinician oversight real rather than nominal; reducing it moves against the requirement. A ignores the explainability trade-off. C misreads the council's objection. D is the opposite of the truth.
  </AccordionItem>

  <AccordionItem title="Q3 · Which TWO conditions make deployment defensible to the clinical council? (Select two)">
    A. Deploy as decision support with a nurse making the final decision and a hard escalation rule on high-harm classes.
    B. Monitor error distribution by presentation type in production.
    C. Attach a disclaimer that the tool is advisory and deploy as-is.
    D. Replace nurses on low-acuity calls to capture the time saving.
    E. Report only aggregate accuracy to the council each quarter.

    **Answer: A and B.** Human oversight with escalation plus monitoring the *distribution* of errors meets the council's real concern. C is blame-shifting, not a safeguard. D removes the oversight that makes it safe. E hides the very risk the council raised.
  </AccordionItem>
</Accordions>

---

## Case 4 · Shadow AI everywhere and a confidentiality clause

**Situation.** Aldergate Advisory is a professional-services firm of about 800 consultants. There is no AI policy. A quiet internal survey finds that 62% of consultants already use consumer AI tools daily – summarising client documents, drafting deliverables, generating code – pasting whatever they are working on into free public chat tools. Productivity anecdotes are glowing.

The problem surfaced when a partner realised that a consultant had pasted a client's confidential strategy document into a free tool whose terms allow the provider to use inputs to improve its models. That client's engagement contract contains an explicit confidentiality clause and a prohibition on disclosing its material to third parties. The firm may already be in breach.

Leadership's instinct splits two ways. The risk partner wants to **ban all AI tools immediately**. The head of consulting wants to **do nothing** because the productivity gains are real and a ban would be ignored anyway. The managing partner asks you, brought in as an AI governance advisor, for a path that neither loses the productivity nor courts the next breach.

**What you are asked.** Recommend a governance response that manages the confidentiality and IP risk without driving the behaviour further underground.

**The evidence you have.**

| Fact | Value (illustrative) |
| --- | --- |
| Consultants | 800 |
| Already using AI daily (survey) | 62% (~496 people) |
| AI policy in place | None |
| Tool used | Free public tool; terms permit training on inputs |
| Trigger event | Confidential client document pasted into a public tool |
| Contractual position | Confidentiality clause; third-party disclosure prohibited |
| Leadership options on the table | Ban all / do nothing |

**Analysis.** Both instincts are wrong for the same reason: they ignore how shadow AI behaves. A blanket ban does not stop use; it stops *visible* use, pushing 496 people onto personal devices where you have no controls and no logs – the opposite of governance (task 1.2.4). Doing nothing accepts an ongoing contractual breach. The exam's answer to shadow AI is neither prohibition nor permissiveness; it is **transparent classification** – a published list of *approved*, *blocked*, and *under-evaluation* tools – paired with a clear acceptable-use rule about *what data* may go into *which* class of tool.

The load-bearing distinction is not "AI good/bad" but **data classification meeting tool classification**. Public client-confidential material must never enter a tool whose terms permit training on inputs; that is the IP and confidentiality risk (task 3.3.3) that caused the incident. The remedy is to provide an *approved*, governed alternative – a tool with terms that do not train on inputs and with access controls – so the sanctioned path is also the easy path. Governance that only says "no" drives shadow AI; governance that says "yes, here, safely" ends it.

Structure follows (task 3.2.1): a small cross-functional owner group – risk, legal, a technology lead, and a respected consultant voice – that maintains the tool list, sets the data rules, and runs a fast exception path so a "no" is rare and a "not yet" is quick. And because the behaviour is already universal, the durable fix is **enablement** (task 4.3.5): tell people plainly what they may and may not do, why the client contract matters, and train them – rather than assume a policy PDF will change habits.

The immediate contractual issue is separate and urgent: legal must assess the possible breach, whether client notification is required, and containment, in parallel with building the governance.

**Recommendation, sequenced.**

1. **Contain the incident now**: legal assesses breach and notification obligations for the affected client; issue an immediate interim rule that no client-confidential material goes into any public tool.
2. **Publish a tool classification** – approved / blocked / under-evaluation – and stand up an approved, governed tool whose terms do not train on inputs.
3. **Write a short acceptable-use rule** keyed to data classification: which data class may enter which tool class; confidential client data only in the approved governed tool.
4. **Form the cross-functional owner group** with a fast exception path so the sanctioned route is the path of least resistance.
5. **Enable the workforce**: communicate the why (the client contract), train on the rules, and make the approved tool genuinely good so shadow use has no reason to persist.
6. **Monitor** adoption of the approved tool as the signal that shadow use is receding.

**What a weaker answer looks like and why.**

- *"Ban all AI tools immediately."* Ends visible use, not use; 496 people move to personal devices and you lose all logs and control – governance in name only.
- *"Do nothing; the productivity is worth it."* Accepts an ongoing contractual breach and guarantees a repeat with a bigger client.
- *"Write a policy PDF and email it round."* A document is not an operating model; without an approved alternative and enablement, habits do not change.
- *"Fine-tune a private model so data stays in-house."* Over-engineered for the actual problem (people need a safe everyday tool now) and does not address the immediate breach or the data-into-public-tools behaviour.

**Which domains and tasks this exercises.** D1 1.2.4 (shadow-AI classification: approved/blocked/under-evaluation); D3 3.2.1 (governance structure with cross-functional accountability), 3.2.3 (access/data-security measures at a strategic level), 3.3.3 (IP and confidentiality risk); D4 4.3.5 (workforce enablement), 4.3.3 (transparent communication).

**Three exam-style questions on this case.**

<Accordions>
  <AccordionItem title="Q1 · With 62% of consultants already using public AI tools, which response best manages shadow-AI risk? (Select one)">
    A. Ban all AI tools firm-wide.
    B. Publish a tool classification (approved/blocked/under-evaluation) with an approved governed alternative and data-use rules.
    C. Take no action because a ban would be ignored.
    D. Require every AI use to be pre-approved individually by the risk partner.

    **Answer: B.** Transparent classification plus a safe sanctioned tool makes the approved path the easy path and ends shadow use. A drives use underground onto personal devices. C accepts the ongoing breach. D creates a bottleneck that will itself be bypassed.
  </AccordionItem>

  <AccordionItem title="Q2 · Why was pasting the client document into the free tool a governance failure specifically? (Select one)">
    A. The tool was slow.
    B. Confidential client data entered a tool whose terms permit training on inputs, breaching a confidentiality clause – a data-classification-meets-tool-classification failure.
    C. The consultant used AI at all.
    D. The output was inaccurate.

    **Answer: B.** The risk is confidential data entering a tool that may train on it, against a contractual prohibition. A is irrelevant. C mislabels all AI use as the problem. D is about output quality, not the confidentiality breach.
  </AccordionItem>

  <AccordionItem title="Q3 · Which TWO steps make the new governance stick rather than becoming a shelved policy? (Select two)">
    A. Provide an approved tool whose terms do not train on inputs and make it genuinely good.
    B. Train consultants on the data rules and explain the client-contract reason.
    C. Publish the policy as a PDF and consider the matter closed.
    D. Rely on individual professionalism with no tooling change.
    E. Audit personal devices without providing an alternative.

    **Answer: A and B.** A safe, good sanctioned tool plus enablement changes behaviour; the approved path must beat the shadow path. C and D change nothing operationally. E is intrusive and still leaves people no safe tool to use.
  </AccordionItem>
</Accordions>

---

## Case 5 · The successful pilot that cannot scale

**Situation.** Trellis Logistics ran a six-month pilot that everyone calls a success. A generative document-processing tool reads carrier paperwork, extracts shipment data, and flags exceptions; in one regional hub it cut manual keying and sped up exception handling. The regional head is delighted and the CEO wants it rolled out to all fourteen hubs by year end.

When you look under the pilot, the readiness picture is thin. The pilot ran on data hand-exported from one hub's system by one analyst who has since moved teams; there is no owner. Each hub stores its paperwork differently, in silos that do not talk to each other. The pilot's cost was small because volume was small, and the tool is billed on a **consumption basis** – per document processed – so cost scales linearly with the fourteen-fold volume increase, with no commitment discount in place. Nobody has defined who owns the tool, the data, or the exceptions once it is enterprise-wide.

The CEO reads the pilot's success as proof the rollout will work. The exam wants you to see the gap between a pilot that proved *feasibility in one place* and an enterprise deployment that needs *readiness everywhere*.

**What you are asked.** Recommend whether to roll out to all fourteen hubs now, and what must be true before scaling.

**The evidence you have.**

| Fact | Value (illustrative) |
| --- | --- |
| Pilot scope | 1 of 14 hubs, 6 months |
| Data source in pilot | Hand-exported by one analyst (now gone) |
| Data across hubs | Siloed, inconsistent formats |
| Owner of tool/data/exceptions | None defined |
| Pricing model | Consumption-based (per document), no commitment discount |
| Pilot monthly cost | $8k at pilot volume |
| Volume increase at full rollout | ~14× |
| CEO's target | All 14 hubs by year end |

**Analysis.** Use the CAF phases as the mental model (task 4.4.1, 4.4.5): the pilot is a **Launch**-phase artefact – it demonstrated incremental value in production in one place. Scaling is the **Scale** phase, and the exam is explicit that the transition from experimental to **production-grade** requires governance and operational foundations the pilot skipped. Assess readiness across the dimensions the guide names (task 4.1.1) and the CAF perspectives (task 4.2.x): the pilot is strong on the *Product* value but weak on *Platform* (data silos), *Operations* (no owner), and *Governance* (no accountability for exceptions).

The two hard gaps are **data** and **ownership** (tasks 4.2.1, 4.2.2). Fourteen silos in inconsistent formats mean the thing that made the pilot easy – one clean hand-export – does not exist at scale. And a hand-export by a person who has left is not a data strategy; enterprise deployment needs defined data ownership and a sharing framework, not heroics.

The cost structure is the trap the CEO will miss (tasks 2.2.5, 2.1.5). At consumption pricing, monthly cost roughly scales with volume: $8k × 14 ≈ **$112k/month**, or ~$1.34M/year, with no floor and no discount. This is where AWS's pricing *structures* matter at a strategic level: a **commitment-based** arrangement (committing to a volume for a discount) can cut predictable, high, steady usage cost materially – but committing before you know the real steady-state volume risks paying for capacity you do not use. The strategic reasoning is: prove the steady-state volume, *then* decide whether a commitment beats consumption. Scaling first and optimising cost never is how a "successful pilot" becomes a budget surprise.

**Recommendation, sequenced.**

1. **Do not roll out to all fourteen hubs now.** The pilot proved feasibility in one place, not readiness everywhere.
2. **Fix the foundations first**: assign an owner for the tool, the data and the exception process; define data ownership and a sharing framework; address the silo and format inconsistency so the tool has a reliable input everywhere.
3. **Model the cost at scale honestly**: ~$112k/month at consumption pricing, and decide the pricing strategy only after observing real multi-hub volume – consumption while volume is uncertain, a commitment once steady-state is known.
4. **Scale in waves, starting with short-term wins** (task 4.4.2): a second and third hub next, each proving the foundations hold, before enterprise-wide rollout.
5. **Set success metrics and a feedback mechanism** (task 4.4.4) so each wave is a gate, not a leap.
6. **Bring the CEO along**: frame it as protecting the pilot's win, not slowing it – an ungoverned fourteen-hub rollout is the fastest way to a public failure.

**What a weaker answer looks like and why.**

- *"Roll out to all fourteen hubs now; the pilot succeeded."* Confuses feasibility-in-one-place with readiness-everywhere; the data silos, missing owner and linear cost are unaddressed.
- *"Sign a large commitment contract to cut the per-document cost before rollout."* Commits to a volume you have not verified; if real volume is lower you pay for unused capacity – optimising cost before proving usage.
- *"Just re-run the same hand-export process in each hub."* Depends on heroics and a person who has left; not a data strategy and not repeatable at fourteen sites.
- *"Appoint an owner and roll out simultaneously."* Names an owner but still ignores the data-silo and cost-at-scale problems; ownership alone does not make the input data usable everywhere.

**Which domains and tasks this exercises.** D4 4.1.1 (readiness dimensions), 4.2.1 (data silos), 4.2.2 (data ownership and sharing), 4.4.1/4.4.5 (phased scaling; experimental to production-grade), 4.4.2 (short-term wins), 4.4.4 (feedback and metrics); D2 2.1.5 (transition considerations), 2.2.5 (cost planning and pricing structure).

**Three exam-style questions on this case.**

<Accordions>
  <AccordionItem title="Q1 · A six-month pilot succeeded in one hub. What does that success actually establish? (Select one)">
    A. That an enterprise rollout to all hubs will work.
    B. Feasibility in one context; enterprise readiness (data, ownership, cost at scale) is a separate question.
    C. That the tool needs no further governance.
    D. That cost will stay proportionally the same.

    **Answer: B.** A pilot proves feasibility in a context; scaling needs data, ownership and cost readiness the pilot did not test. A over-generalises. C ignores the production-grade governance the transition requires. D is false under consumption pricing, where cost scales with volume.
  </AccordionItem>

  <AccordionItem title="Q2 · The tool is billed per document with no commitment discount. What is the sound cost strategy for scaling? (Select one)">
    A. Sign a large volume commitment before rollout to lock in a discount.
    B. Stay on consumption pricing while volume is uncertain, then evaluate a commitment once steady-state volume is known.
    C. Ignore cost; the pilot was cheap.
    D. Switch to a seat-based tool regardless of fit.

    **Answer: B.** Prove the real volume, then decide whether a commitment beats consumption; committing early risks paying for unused capacity. A commits before knowing volume. C ignores that consumption cost scales ~14×. D changes pricing model without regard to fit.
  </AccordionItem>

  <AccordionItem title="Q3 · Which TWO foundations must be fixed before scaling beyond one hub? (Select two)">
    A. Define data ownership and a sharing framework across hubs.
    B. Assign an owner for the tool and the exception process.
    C. Increase the model's context window.
    D. Re-run the analyst's manual export in each hub.
    E. Announce the full rollout date to build momentum.

    **Answer: A and B.** Data ownership/sharing and a named owner are the readiness gaps that make scale safe; the pilot skipped both. C is a technical detail out of scope for the readiness gap. D depends on heroics and does not scale. E commits to a date before the foundations exist.
  </AccordionItem>
</Accordions>

---

## Case 6 · Adoption stalled at 11% and the technology works

**Situation.** Fenwick Insurance is eighteen months into an AI transformation. It bought enterprise licences for a generative assistant for 3,000 knowledge workers, ran a launch event, and published a training portal. The technology works well; a proof-of-concept in claims showed real time savings, and the tool is stable and well-reviewed by the people who use it.

The problem is that almost nobody uses it. Telemetry shows **11% of licensed users** are active in a given month, and active use is concentrated in two enthusiastic teams. The rest tried it once or never logged in. The CFO is asking why the firm is paying for 3,000 seats to get 330 users. The transformation sponsor's instinct is to buy more training and send another all-staff email.

You interview across the business. The pattern is not a technology problem. Managers do not use the tool themselves and do not ask their teams to; some quietly worry it makes their expertise look replaceable. Frontline staff are unsure whether using it is encouraged or a shortcut that will be held against them, and several fear it is a prelude to job cuts – a fear nobody has addressed. The two teams with high adoption both have a manager who models the behaviour and a couple of informal champions.

**What you are asked.** Recommend how to move adoption, given that the technology is not the constraint.

**The evidence you have.**

| Fact | Value (illustrative) |
| --- | --- |
| Licensed seats | 3,000 |
| Monthly active users | 11% (~330) |
| Where adoption is high | 2 teams with a modelling manager + champions |
| Technology quality | Good; stable; POC showed real savings |
| Training provided | Portal + launch event |
| Manager behaviour | Mostly not using it, not asking teams to |
| Unaddressed fear | Job security / role impact |
| Seat pricing | Seat-based (paying per licence regardless of use) |

**Analysis.** The exam's change-management material (task 4.3) is built for exactly this: the technology works, the value is proven, and adoption still stalls because the *human* system was not led. More training and another email treat a leadership and cultural problem as an information problem – and the evidence says information is not the gap, because the two successful teams did not get more training; they had a manager who modelled the behaviour and champions who normalised it.

Diagnose the barriers by name (task 4.3.4): **manager behaviour** (leaders who do not use it signal it is optional), **fear about role impact** (unaddressed job-security anxiety suppresses use), and the **absence of champions** outside the two teams. The seat-based cost structure sharpens the urgency – the firm pays for 3,000 seats whether or not they are used, so at, illustratively, $100/seat/month the annual licence bill is 3,000 × $100 × 12 = $3.6M/year, of which ~89% (≈$3.2M) is idle at 11% adoption – but cost is the symptom, not the lever.

The levers are behavioural. **Executive sponsorship and manager modelling** (task 4.3.1): make managers use it and expect their teams to, because the successful teams show that is the difference. **Champions** (task 4.3.1): recruit and empower the informal enthusiasts as a network, not a one-off. **Transparent communication about role impact** (task 4.3.3): address the job-security fear directly and honestly – what the tool is for, how roles change toward oversight and higher-value work (task 4.3.6), and what will not happen – because unspoken fear is a stronger brake than any missing feature. And **measure leading indicators** (task 2.2.4, 4.4.4): weekly active use by team, manager usage, champion activity – signals that predict adoption – rather than waiting for a lagging productivity number.

**Recommendation, sequenced.**

1. **Stop the reflex spend** on more generic training and mass emails; they address a gap that is not the constraint.
2. **Enlist managers as the primary lever**: make usage a visible leadership expectation, and have managers model it in their own work – replicate what the two successful teams did.
3. **Build a champion network** from the existing enthusiasts, resourced and recognised, seeded into every team.
4. **Address role-impact fear head-on** with honest, transparent communication about how roles shift toward oversight and higher-value work, and what job impact will and will not occur.
5. **Track leading indicators** (weekly active use, manager usage, champion activity) by team and act on the laggards.
6. **Only then** revisit seat counts – right-size licences to real, growing usage once the behavioural levers are working, rather than cutting seats and killing momentum.

**What a weaker answer looks like and why.**

- *"Buy more training and send another all-staff email."* Treats a leadership and cultural problem as an information gap; the successful teams prove information was never the constraint.
- *"Cut licences to 330 to stop wasting money."* Optimises the symptom and abandons the transformation; it saves cost by giving up on 90% of the intended value.
- *"Mandate usage and track it punitively."* A mandate without addressing the job-security fear deepens resistance and produces gaming, not genuine adoption.
- *"Wait for the lagging productivity metric before acting."* Leading indicators exist precisely so you do not wait; by the time the lagging number moves, the licence spend is long wasted.

**Which domains and tasks this exercises.** D4 4.3.1 (sponsorship, manager modelling, champions), 4.3.3 (communication about role impact), 4.3.4 (cultural barriers, leadership intervention), 4.3.6 (transition roles toward oversight); D2 2.2.4 (leading indicators), 2.2.1 (intangible benefits like productivity and satisfaction); D4 4.4.4 (feedback and success metrics).

**Three exam-style questions on this case.**

<Accordions>
  <AccordionItem title="Q1 · The technology works but only 11% of licensed users are active. What is the BEST first response? (Select one)">
    A. Buy more training and send another all-staff email.
    B. Enlist managers to model usage and expect it of their teams, replicating the two teams that already adopted.
    C. Cut licences to match current usage.
    D. Mandate usage and monitor it punitively.

    **Answer: B.** The successful teams show manager modelling is the difference; adoption here is a leadership and cultural problem, not an information one. A repeats an information fix that the evidence says is not the gap. C abandons the transformation. D deepens fear and produces gaming.
  </AccordionItem>

  <AccordionItem title="Q2 · Interviews reveal staff fear the tool is a prelude to job cuts. How should a leader address this? (Select one)">
    A. Avoid the topic so as not to alarm people.
    B. Communicate transparently about how roles shift toward oversight and higher-value work, and what will and will not happen.
    C. Promise there will never be any change to any role.
    D. Leave it to each manager to handle privately.

    **Answer: B.** Transparent communication about role impact is a named change-management task; unaddressed fear suppresses adoption. A lets fear fester. C makes a promise that is not credible and will backfire. D leaves the hardest, firm-wide message to inconsistent local handling.
  </AccordionItem>

  <AccordionItem title="Q3 · Which TWO signals are the right leading indicators of adoption to track and act on? (Select two)">
    A. Weekly active users by team and manager usage rates.
    B. Champion activity across teams.
    C. The lagging annual productivity figure only.
    D. Total seats purchased.
    E. Number of all-staff emails sent.

    **Answer: A and B.** Weekly active use, manager usage and champion activity are leading indicators that predict adoption and can be acted on early. C is a lagging metric you should not wait for. D is a cost input, not an adoption signal. E measures activity, not outcome.
  </AccordionItem>
</Accordions>

---

## How to use these cases

Work each case before reading its analysis: write your recommendation, then compare. The skill the exam rewards is not recalling a definition – it is holding several criteria at once, resisting the sponsor's fixed belief, doing the arithmetic, and sequencing the fix so foundations come before scale and governance comes before launch. When you can reconstruct the reasoning trace, not just recognise the answer, you are ready for the long items. Return to the relevant domain page for depth: [D1](/aws/aib-c01/domains/d1-ai-fundamentals-and-literacy/), [D2](/aws/aib-c01/domains/d2-ai-strategy-and-business-value/), [D3](/aws/aib-c01/domains/d3-governance-and-responsible-ai/), [D4](/aws/aib-c01/domains/d4-readiness-leadership-transformation/), then sit the [mock exams](/aws/aib-c01/practice-exam/).
