AI Cert Prep
Type to search documentation.

Appendix · AWS

KPI and Metrics Library

A catalogue of measurable AI KPIs by business function with formulas, baselines and gaming risks, plus tangible vs intangible benefits, leading vs lagging pairs, adoption metrics, a KPI tree, baseline-construction methods and vanity metrics to avoid — for AIB-C01.

This is a working catalogue for the measurement work Domain 2 tests (tasks 2.2.1–2.2.4): KPIs by business function, tangible versus intangible benefits, leading versus lagging indicators, adoption and health metrics, a KPI tree, and how to build a baseline when none exists. It is independent preparation. The AIB-C01 beta exam does not ask you to instrument anything; it asks you to pick the right metric, define a defensible baseline, and spot a metric that measures activity instead of value. Every formula here stays at that decision level.

Baseline before deployment

The single most-repeated Domain 2 discriminator: value you did not baseline is value you cannot claim. When a stem says a team wants to prove ROI after launch and never measured the before-state, the correct answer reconstructs a baseline (§9) — an after-only number is not evidence.

How the metrics fit together

text
BUSINESS OBJECTIVE (e.g. cut cost-to-serve 15%)
│
▼
VALUE KPI (lagging) ── cost per contact, ROI, revenue
│
▼
OPERATIONAL KPI ──── handle time, containment, rework rate
│
▼
ADOPTION / HEALTH (leading) ── activation, weekly active use,
│ task completion, quality score
▼
INSTRUMENTED SIGNAL ── events you can actually count

Read it top-down to design, bottom-up to explain. Executives care about the objective; the leading indicators near the bottom are what tell you three weeks in whether the lagging value at the top will ever arrive.

1. KPIs by business function

Each KPI below carries a definition, a formula, where the baseline comes from, how to set a target, and the way it gets gamed. Pick the smallest set that ties to the objective; more metrics is not more insight.

KPIDefinition / formulaBaseline sourceTarget guidanceGaming risk
Containment rateContacts resolved without a human / total contactsCurrent self-service logsSet below 100%; over-containment traps customersDeflecting hard cases into a dead end to boost the ratio
Average handle time (AHT)Total handle time / contacts handledHistorical CRM timestampsRealistic minutes saved, not zeroRushing to close, hurting resolution quality
First-contact resolutionResolved on first contact / totalQA samplingImprove, not maximise at cost of accuracyMarking reopened issues as new
CSAT / CESSurvey score after AI-handled contactPre-AI survey meanMatch or beat human baselineCherry-picking who gets surveyed

Every table shares one rule: the baseline source is a before number, and every KPI has a gaming risk. A KPI with no gaming risk column is a KPI you have not thought hard enough about.

2. Tangible vs intangible benefits

Task 2.2.1 splits benefits explicitly. Tangible benefits convert to currency directly; intangible ones do not, which is why they get dropped from business cases — and why the mistake is not to make them defensible.

TypeExamplesHow to defend it
TangibleCost reduction, revenue growth, hours saved × loaded rate, error-cost avoidedTie to a GL line or a headcount rate; show before and after
IntangibleCustomer satisfaction, employee productivity, brand trust, decision speed, risk reductionProxy it: CSAT delta, retention lift, survey-measured time saved, incidents avoided × expected cost

The technique for intangibles is to attach a measurable proxy and, where possible, a conservative monetary bridge. “Employees are happier” is not defensible; “eNPS rose 8 points and regretted attrition fell 3 points, worth roughly one avoided backfill per quarter at £X” is. Always mark such figures as indicative.

IntangibleDefensible proxyMonetary bridge (conservative)
Employee productivitySelf-reported + sampled hours saved per weekHours × loaded hourly rate × active users
Customer satisfactionCSAT / CES / NPS delta on AI-handled interactionsRetention lift × customer lifetime value
Decision quality/speedTime-to-decision; rework rateValue of faster cycle; cost of avoided rework
Risk reductionIncidents avoided vs baseline rateExpected loss avoided × probability

3. Leading vs lagging indicators

Task 2.2.4 asks you to identify leading indicators that predict success. Lagging indicators confirm value after the fact; leading indicators tell you early whether it is coming. Pair them.

Programme goalLeading indicator (early, predictive)Lagging indicator (confirms value)
Support cost reductionWeekly active agents using the assistant; task-completion rateCost per contact; AHT; CSAT
Sales liftReps adopting the tool; assisted opportunities createdConversion rate; revenue per rep
Developer velocitySuggestions accepted; PRs assistedCycle time; change-failure rate
Content programmeDrafts started with AI; review pass ratePublish cycle time; campaign ROI

If a leading indicator is flat three weeks in, the lagging value will not arrive — that is the point of watching it. A programme that reports only lagging metrics finds out too late.

4. Adoption and health metrics

Value requires use. These metrics tell you whether the tool is actually being adopted well, and they are the leading indicators most programmes forget to instrument.

MetricDefinitionWhat it warns you about
ActivationUsers who reached first successful use / provisionedOnboarding friction; licences bought but unused
Weekly active useDistinct users using it in a rolling week / target usersPilot enthusiasm fading; no habit forming
Task completion rateTasks finished with AI / tasks startedThe tool fails on real work, not demos
Containment rateHandled without human handoff / totalOver- or under-automation
Escalation rateHandoffs to a human / totalWhere the tool hits its limits
Rework rateOutputs needing correction / totalQuality problem masquerading as productivity
Quality sampling scoreHuman-graded sample against a rubricThe number CSAT can hide

Containment and escalation are two sides of one coin: read them together, because a high containment rate with a rising rework rate means you are trapping users, not helping them.

5. A worked KPI tree

Start from the objective and decompose until you reach something you can count. This is the artefact that connects an executive goal to an instrumented signal, and it is the shape the exam rewards.

text
OBJECTIVE: Reduce cost-to-serve in support by 15% in 12 months
│
├── Value KPI: Cost per contact (lagging)
│ │
│ ├── Driver: Contacts handled without a human → Containment rate
│ ├── Driver: Time per human-handled contact → AHT
│ └── Guardrail: Customer satisfaction → CSAT (must not fall)
│
└── Leading indicators (weeks 1–6):
├── Weekly active agents using the assistant
├── Task-completion rate on real tickets
└── Escalation rate trend (falling = tool coping)

Worked arithmetic: 500,000 contacts/year at £6.00 each = £3.0m. A 30% containment rate at £0.40 per contained contact, with the remaining 70% at a reduced £5.40 AHT-adjusted cost, gives 500,000 × (0.30 × 0.40 + 0.70 × 5.40) = 500,000 × (0.12 + 3.78) = £1.95m — a 35% cost reduction if CSAT holds. The guardrail metric is why the tree includes CSAT: a cost win that drops satisfaction is not a win.

6. Baselines: constructing one when none exists

Task 2.2.2 requires a baseline before implementation, but the common real-world blocker is that no baseline was ever recorded. There are four defensible ways to build one; pick by data availability and cost.

MethodHowWhen to useWeakness
Time-and-motion sampleObserve/measure a representative sample of the current processNo historical logs existSampling bias; observer effect
Historical proxyUse existing system logs (CRM, ERP, ATS) as the before-stateTimestamps/records already capturedProxy may not match the exact metric
Control groupRun AI for one group, hold another unchanged, compareYou can split fairly and ethicallyContamination; group comparability
Staged rolloutCompare cohorts before and after phased enablementBig-bang launch is risky anywayTime trends confound the comparison

The exam’s preferred answer to “we have no baseline” is build one now by the cheapest credible method — never “measure after launch and assume the difference is the AI”. A control group or staged rollout is the strongest because it isolates the AI effect from background change.

7. Metrics that look good and mean nothing

Vanity metrics feel like progress and predict nothing. Recognising them is a direct exam skill.

Vanity metricWhy it is temptingReplace with
Number of prompts / queries runBig number, easy to pullTask-completion rate; value KPI
Licences purchasedLooks like adoptionWeekly active use; activation
Model accuracy in isolationTechnical and impressiveBusiness outcome the accuracy serves
Total outputs generatedShows the tool is “busy”Outputs that passed review; rework rate
Pilot NPS from volunteersEnthusiastic early adoptersSampled quality score across all users

The tell of a vanity metric: it can rise while the business outcome is flat or falling. If a metric can go up while cost-per-contact and CSAT do nothing, it is measuring activity, not value.

Key takeaways

  • Choose the smallest KPI set that ties to the objective; every KPI needs a formula, a before-baseline and an acknowledged gaming risk.
  • Make intangible benefits defensible with a measurable proxy and a conservative, clearly-labelled monetary bridge — do not drop them.
  • Pair a leading indicator (adoption, task completion) with each lagging value KPI so you learn early, not late.
  • Read containment against escalation and rework; a high containment rate with rising rework means you are trapping users.
  • Build the KPI tree from objective to instrumented signal, and keep a guardrail metric (like CSAT) so a cost win cannot quietly destroy quality.
  • With no baseline, construct one now — time-and-motion, historical proxy, control group or staged rollout — never measure after launch and assume the delta is the AI.
  • A metric that can rise while the business outcome is flat is a vanity metric; replace it.

Last updated Sep 18, 2026