# D1 · Scoping AI Solutions

Framing a problem for AI, eliciting requirements, defining success metrics and a cost envelope, deciding build versus buy, judging feasibility and writing a solution plan.

import { Accordions, AccordionItem } from '@prosefly/astro-components';

This domain is about **12%** of the OAI-API mock – roughly **7 of 60 items** – and it mirrors the Academy *Scope AI Solutions* course (30 min). It tests the work you do *before* a single API call: turning a vague business ask into a defined job, a measurable success criterion and a cost envelope, and deciding whether AI is even the right tool. Most items describe a stakeholder request and ask what you should establish first, or whether the problem is a good fit for a language model at all.

## What you need to know

Scoping is requirement elicitation for probabilistic software. You clarify the *job to be done*, the *inputs and outputs*, the *quality bar* and *how it will be measured*, the *volume and latency* the system must sustain, and the *cost and risk envelope* the business will accept. You decide build-vs-buy (a Responses API integration, a hosted product, or no AI at all), and you judge feasibility honestly – some tasks a model does well, some it does badly, and some it cannot verify. The output of scoping is a short **solution plan** that names the metric you will move and the smallest thing you can ship to test it.

## Learning objectives

By the end of this page you should be able to:

1. **Reframe** a business request as a well-formed AI task with defined inputs, outputs and constraints.
2. **Elicit requirements** – volume, latency, quality bar, data sensitivity, integration points.
3. **Define a success metric** that is measurable before you build.
4. **Estimate a cost envelope** from token volume and model price, and check it against the business value.
5. **Judge feasibility** and choose build-vs-buy, including "do not use AI here".
6. **Write a solution plan** that scopes the smallest useful first version.

---

## 1.1 Reframing a request as a task

Stakeholders describe outcomes ("make support faster"); you need a task ("draft a first-response reply from the ticket text and the last three tickets from the same customer"). A well-formed task names the inputs the model will see, the output shape it must produce, and the boundary of what it must *not* do.

| Vague request | Reframed task | Inputs | Output |
| --- | --- | --- | --- |
| "Summarise our meetings" | Produce a 5-bullet action list with owners from a transcript | Transcript text | Markdown list, one owner per action |
| "Answer policy questions" | Answer HR policy questions grounded only in the policy PDFs, with citations | Question + retrieved policy chunks | Answer + cited sources |
| "Triage tickets" | Classify a ticket into one of 8 queues and a priority | Ticket subject + body | JSON `{queue, priority}` |
| "Help engineers" | Suggest a code fix for a failing test, for human review | Repo + failing test output | Patch proposal |

:::tip[Assessment signal]
Stems that say "a stakeholder asks you to…", "the business wants…", or "before you start building" are scoping items. The correct answer usually **clarifies or measures something first** rather than jumping to a model or a prompt.
:::

## 1.2 Eliciting requirements

Six questions decide almost every design choice downstream. Ask them before you pick a model.

| Requirement | Question | Why it drives design |
| --- | --- | --- |
| Volume | How many requests per day, and peak per minute? | Sets rate limits, batch vs realtime, cost |
| Latency | Interactive (user waiting) or background? | Streaming, background mode, `fast` vs `flex` |
| Quality bar | What is "good enough", and who decides? | Model choice, reasoning effort, review gate |
| Data sensitivity | PII, regulated data, residency constraints? | Retention, `store`, residency, network controls |
| Inputs | Text, images, files, structured data? | Model capability, file inputs, RAG |
| Integration | Where does the output go, in what format? | Structured outputs, function calling |

## 1.3 Defining a success metric

If you cannot measure it, you cannot improve it and you cannot prove it works. A scoping metric must be defined *before* build and be evaluable on a held-out dataset (this is where D1 hands off to D3, evals).

```text
Task type            Good success metric                Bad metric
─────────────────    ───────────────────────────────    ─────────────────
Classification       accuracy / F1 on a labelled set    "feels accurate"
Extraction           field-level exact match            "looks complete"
Grounded Q&A         % answers supported by a citation  "sounds right"
Summarisation        rubric score by an LLM grader      "reads well"
Support drafting     % accepted by agent without edit   "faster" (unmeasured)
```

:::tip[Assessment signal]
When an option says "we will know it is working when customers are happier" and another says "we will measure the percentage of drafts an agent accepts unedited on a 200-ticket sample", the measurable, pre-defined metric is correct. Vague outcomes are distractors.
:::

## 1.4 The cost envelope

Scope the cost before you build; a solution that works but costs more than the value it creates is a failed scope. Estimate from expected tokens and the model price. The September 2026 prices per **million tokens**:

| Model | Input $/MTok | Output $/MTok |
| --- | --- | --- |
| `gpt-6-astra` | $10 | $50 |
| `gpt-5.6-sol` | $4 | $20 |
| `gpt-5.6-terra` | $2 | $12 |
| `gpt-5.6-luna` | $0.20 | $1.20 |

Worked example: a triage classifier on `gpt-5.6-luna`, 800 input tokens and 40 output tokens per ticket, 50,000 tickets/day.

```text
Input:  50,000 × 800  = 40,000,000 tok/day = 40 MTok × $0.20 = $8.00/day
Output: 50,000 × 40   =  2,000,000 tok/day =  2 MTok × $1.20 = $2.40/day
Daily total ≈ $10.40   →  ≈ $312/month
```

The same task on `gpt-5.6-sol` would be roughly `40 × $4 + 2 × $20 = $200/day` (~$6,000/month) for no accuracy gain on a task Luna handles – the cost envelope is what rejects that choice.

## 1.5 Build versus buy versus don't

| Option | Choose when | Watch out for |
| --- | --- | --- |
| **Build on the API** | You need control, custom data, or the task is core to your product | You own evals, ops, cost and safety |
| **Buy a hosted product** | A commodity capability exists (transcription, generic chat) and speed-to-value wins | Data flow, lock-in, per-seat cost at scale |
| **Use ChatGPT / no code** | A knowledge worker can do it in the UI with Projects and files | Not repeatable or auditable at volume |
| **Don't use AI** | The task needs a guaranteed-correct, deterministic answer | Forcing AI onto arithmetic, lookups, or exact rules |

:::caution[The "AI can't verify itself" test]
If the task requires an answer that must be *provably* correct and there is no cheap external check (regulatory lookups, financial reconciliation with a ground truth, safety-critical instructions), a probabilistic model is the wrong core – use it to assist a deterministic system, not to replace it.
:::

## 1.6 Judging feasibility

```text
             Does a ground-truth or a cheap check exist?
                        │
         ┌──────────────┴───────────────┐
         │ Yes                          │ No
         ▼                              ▼
   Measurable → good AI fit.      Can a human review each output
   Build with an eval loop.       at the required volume?
                                   │
                        ┌──────────┴──────────┐
                        │ Yes                 │ No
                        ▼                     ▼
                  Human-in-the-loop      Reconsider scope: narrow
                  assist is feasible.    the task, add a check, or
                                         don't ship AI here.
```

## 1.7 The solution plan template

The deliverable of scoping is a one-page plan. Every item on the mock that asks "what should the plan include" expects these fields.

```text
Problem:        the business pain, in one sentence
Job to be done: the specific task, with inputs and outputs
Users & volume: who, how many requests/day, peak/min
Success metric: measurable, on a held-out dataset, target value
Constraints:    latency, data sensitivity, residency, budget
Cost envelope:  $/request estimate × volume vs. value created
Approach:       build/buy/no-AI, model, retrieval? tools?
Risks & gates:  failure modes, human review points
First version:  smallest shippable slice + how you'll evaluate it
```

## Decision framework

Use the **SCOPE** framework to turn any request into a plan you could act on tomorrow.

| Letter | Step | Question to answer | Output |
| --- | --- | --- | --- |
| **S** | State the job | What exact task, with what inputs and outputs? | One-sentence task definition |
| **C** | Criteria | How will we measure "good enough"? | A metric and a target on a dataset |
| **O** | Operating limits | Volume, latency, data sensitivity, budget? | Constraint list + cost envelope |
| **P** | Path | Build, buy, or no AI? Which model and tools? | Approach and model choice |
| **E** | Experiment | What is the smallest version that tests the metric? | First-version scope + eval plan |

The most valuable habit the course teaches is finishing **C** before starting **P**: teams that pick a model before defining the metric almost always over-buy.

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Picking the model first | It feels like progress; models are exciting | Define the metric and cost envelope, then pick the cheapest model that meets them |
| No measurable success metric | "It looks good" is easy; a dataset is work | Write a metric evaluable on a held-out set before building |
| Scoping the ideal system, not the first version | Ambition and stakeholder pressure | Scope the smallest slice that tests the metric |
| Ignoring volume and peak | Demos run once; production runs constantly | Estimate requests/day and peak/min; size limits and cost from them |
| Forcing AI onto deterministic work | AI is the mandate of the quarter | Use deterministic code for exact rules; let AI assist, not decide |
| Skipping data-sensitivity questions | It surfaces late, as a launch blocker | Ask about PII, regulated data and residency during elicitation |
| Treating cost as an afterthought | Token math is tedious | Compute $/request × volume and compare to value in the plan |
| Confusing "faster" with a metric | Speed is intuitive but unmeasured | Define what faster means numerically (e.g., median handle time) |

## Scenario challenge

**Scenario.** You are a developer at a logistics company. The VP of Operations says: "Our dispatchers waste hours reading driver incident reports. Build an AI that handles them." There is no dataset, no defined output, and the VP expects a demo in a week. Incident reports are free text, sometimes contain injury details (regulated), arrive at roughly 3,000/day with spikes after storms, and currently a dispatcher reads each one and either files it, escalates to safety, or requests more detail.

**Expert reasoning trace.**

1. **Reframe the request into a job.** "Handle them" is an outcome, not a task. The real job is: given a report, classify it into `file / escalate / need-more-detail` and extract structured fields (date, location, severity, injury flag). That is a classification-plus-extraction task, not open-ended chat.
2. **Define the metric before building.** Ask the safety team for 200 historical reports with the decision a dispatcher actually made. That labelled set gives an accuracy target and, critically, a way to check the *escalation* decision – the high-stakes one where a false negative is dangerous.
3. **Surface data sensitivity early.** Injury details are regulated. That drives retention (`store: false` or short retention), residency questions, and a mandatory human review gate on any `escalate` decision – you never auto-close a safety-relevant report.
4. **Estimate the cost envelope.** 3,000/day at ~1,200 input + 60 output tokens on `gpt-5.6-luna` is ~`3.6 MTok × $0.20 + 0.18 MTok × $1.20 ≈ $0.72 + $0.22 = $0.94/day`. Trivial versus dispatcher hours saved – the envelope is not the constraint here; correctness on escalation is.
5. **Scope the first version, not the ideal one.** Ship the classifier with human confirmation on every `escalate`, measured against the 200-report set. Defer auto-filing until the eval proves the escalation recall is high enough.
6. **Reject the demo trap.** A week-one demo that classifies with no eval and no review gate would look impressive and be dangerous. The plan says: build the eval set first, then the classifier, then measure.

**Exam-correct decision:** establish the labelled dataset and the escalation metric first, scope a human-in-the-loop first version, and size cost from real volume. **Not** "start prompting `gpt-6-astra` and demo it", **not** "auto-close low-severity reports before measuring recall".

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "Pick `gpt-6-astra` so quality is never the problem" | Best model feels safest | Scope the metric and cost first; the cheapest model that passes the eval wins |
| "Ship the demo this week to show progress" | Stakeholder pressure | A demo without a metric or a review gate is scope failure, not progress |
| "The metric is that users are happier" | Sounds customer-focused | A metric must be measurable on a held-out dataset before build |
| "Automate the whole workflow end to end" | Ambition | Scope the smallest slice that tests the metric; keep humans on high-stakes steps |
| "AI can do anything, so it fits" | Hype | If the task needs provable correctness with no cheap check, AI is the wrong core |
| "Cost doesn't matter, tokens are cheap" | Individually true | At volume, $/request × requests/day is the number that kills or saves a project |

## Practice questions

Each item states how many responses to select. Commit before revealing.

<Accordions>
  <AccordionItem title="Q1 · A product manager asks you to 'use AI to improve onboarding'. What should you establish FIRST? (Select one)">
    A. Which model has the largest context window.
    B. The specific task, its inputs and outputs, and a measurable success metric.
    C. Whether to use the Agents SDK or the Agents API.
    D. The prompt wording for the first version.

    **Answer: B.** Scoping starts by reframing an outcome into a defined, measurable task. Model context (A), runtime (C) and prompt wording (D) are all downstream `P`-step choices that depend on the job and metric you have not defined yet.
  </AccordionItem>

  <AccordionItem title="Q2 · Which of the following is a well-formed success metric for a ticket-triage classifier? (Select one)">
    A. Customers report being happier with support.
    B. The model feels accurate in testing.
    C. Classification accuracy of at least 92% on a held-out set of 500 labelled tickets.
    D. Responses are generated in under two seconds.

    **Answer: C.** A metric must be measurable on a held-out dataset with a target. Happiness (A) is an unmeasured outcome, "feels accurate" (B) is subjective, and latency (D) is a constraint, not a quality metric for classification.
  </AccordionItem>

  <AccordionItem title="Q3 · A task requires returning a customer's exact contractual renewal date, which exists in a structured billing database. What is the BEST approach? (Select one)">
    A. Ask `gpt-6-astra` at `max` reasoning effort to recall the date.
    B. Have the model call a function that queries the billing database and return the exact value.
    C. Fine-tune a model on all past renewal dates.
    D. Put the entire billing table in the prompt every request.

    **Answer: B.** Exact, provable facts belong in a deterministic lookup the model calls as a tool, not in model recall. Recall (A) can hallucinate, fine-tuning (C) bakes in stale data, and dumping the whole table (D) is expensive and still not authoritative.
  </AccordionItem>

  <AccordionItem title="Q4 · You are scoping a summarisation feature for 20,000 documents per day, ~1,000 input and ~150 output tokens each, quality bar met by `gpt-5.6-luna`. Roughly what is the daily cost? (Select one)">
    A. About $0.76.
    B. About $7.60.
    C. About $76.
    D. About $760.

    **Answer: B.** Input: `20,000 × 1,000 = 20 MTok × $0.20 = $4.00`. Output: `20,000 × 150 = 3 MTok × $1.20 = $3.60`. Total ≈ `$7.60/day`. The discriminator is doing the `tokens ÷ 1,000,000 × price/MTok` arithmetic correctly: option A is 10× too low, and C and D are 10× and 100× too high.
  </AccordionItem>

  <AccordionItem title="Q5 · A stakeholder wants an AI that computes exact monthly financial reconciliations that must always match the ledger. Which TWO statements should shape your scope? (Select two)">
    A. Exact reconciliation is deterministic work; the model should orchestrate deterministic calculations, not perform the arithmetic itself.
    B. Any output touching financial figures needs a verification step against ground truth before use.
    C. `gpt-6-astra` at `max` effort guarantees correct arithmetic.
    D. Fine-tuning on past reconciliations makes the arithmetic exact.
    E. Because it is internal, no verification is needed.

    **Answer: A and B.** Exact arithmetic is deterministic and must be computed by code (or a tool) with a verification step against the ledger; the model coordinates. No reasoning effort (C) or fine-tuning (D) makes a probabilistic model provably exact, and "internal" (E) does not remove the need to verify financial figures.
  </AccordionItem>

  <AccordionItem title="Q6 · Your VP wants a full autonomous incident-handling system in one week, with no labelled data available. What is the MOST appropriate first step? (Select one)">
    A. Build the full autonomous system and demo it.
    B. Assemble a small labelled dataset of past incidents and their correct dispositions, then scope a human-in-the-loop first version measured against it.
    C. Pick the largest model and enable every tool.
    D. Skip the metric and iterate on the prompt until the demo looks good.

    **Answer: B.** With no data and a high-stakes decision, you must build the evaluation dataset and scope a reviewed first version. Building everything (A) or maximising tools (C) skips measurement, and prompt-tuning to a good-looking demo (D) has no metric behind it.
  </AccordionItem>

  <AccordionItem title="Q7 · A commodity need – transcribing support calls to text – has a mature hosted product available. Your team is small and speed-to-value matters. Which factor MOST justifies buying over building? (Select one)">
    A. Building would let you claim you built it.
    B. Transcription is a solved commodity; buying delivers value faster and lets your team focus on the differentiated part.
    C. Hosted products are always cheaper at every scale.
    D. Building avoids all data-flow considerations.

    **Answer: B.** Buy commodity capability, build the differentiator. Bragging rights (A) are not a business reason, hosted is not always cheaper at scale (C), and buying introduces, not removes, data-flow considerations (D).
  </AccordionItem>

  <AccordionItem title="Q8 · Which requirements are ESSENTIAL to elicit during scoping because they change the architecture? (Select two)">
    A. The daily request volume and peak per minute.
    B. Whether inputs contain PII or regulated data.
    C. The founder's favourite model.
    D. The colour of the dashboard.
    E. Whether the office uses Mac or Windows.

    **Answer: A and B.** Volume/peak sizes rate limits, batch-vs-realtime and cost; data sensitivity drives retention, residency and review gates – both change the architecture. Model preference (C), UI colour (D) and OS (E) do not.
  </AccordionItem>

  <AccordionItem title="Q9 · A feature must respond while a user waits in a chat UI, and another must process a nightly backlog of 2M documents. What does correct scoping conclude? (Select one)">
    A. Both should use realtime streaming for consistency.
    B. The interactive feature needs low latency (streaming, possibly fast mode); the backlog is a good fit for Batch, which trades latency for lower cost.
    C. Both should use the Batch API.
    D. Both should use background mode with webhooks.

    **Answer: B.** Latency requirements differ: interactive work needs streaming/low latency, while a large non-urgent backlog fits Batch's cheaper, higher-latency processing. Forcing one mode on both (A, C) ignores the latency requirement; background mode (D) suits the backlog but not the waiting user.
  </AccordionItem>

  <AccordionItem title="Q10 · Which item belongs in the 'Success metric' line of a solution plan? (Select one)">
    A. The model ID you will use.
    B. A target value on a held-out dataset, e.g. 'field-level exact match ≥ 95% on 300 labelled invoices'.
    C. The list of engineers on the project.
    D. The prompt template.

    **Answer: B.** The success-metric line names a measurable target on a dataset. Model ID (A) is the approach line, staffing (C) is not a metric, and the prompt (D) is an implementation detail, not a success criterion.
  </AccordionItem>

  <AccordionItem title="Q11 · A director insists on `gpt-6-astra` for a high-volume extraction task that `gpt-5.6-luna` handles at 96% on your eval. What is the BEST response? (Select one)">
    A. Comply; the best model is always safest.
    B. Show the eval and the cost envelope: Luna meets the metric at roughly a fiftieth of the input price, so Astra adds cost without measurable benefit.
    C. Use Astra but at `low` effort to save money.
    D. Skip the eval and split traffic between both.

    **Answer: B.** The scoping discipline is 'the cheapest model that passes the eval', backed by the cost math. Complying by default (A) over-buys, Astra-at-low (C) still costs far more per token than Luna, and splitting traffic without an eval (D) has no basis for the decision.
  </AccordionItem>

  <AccordionItem title="Q12 · During scoping, you discover the task's outputs feed an irreversible action (auto-cancelling shipments). How should the plan reflect this? (Select one)">
    A. Nothing changes; the model is accurate enough.
    B. Add a mandatory human review gate before the irreversible action and define a metric specifically for the decision that triggers it.
    C. Increase reasoning effort to `max` and skip review.
    D. Log the action after the fact for auditing only.

    **Answer: B.** Irreversible actions require a review gate and a targeted metric on the triggering decision. Trusting accuracy (A) or raising effort (C) does not make an irreversible auto-action safe, and after-the-fact logging (D) does not prevent the harm.
  </AccordionItem>
</Accordions>

## Key takeaways

- Scoping turns an outcome into a defined task with inputs, outputs and a measurable metric.
- Define the success metric on a held-out dataset **before** choosing a model.
- Elicit volume, latency, quality bar, data sensitivity, inputs and integration – they drive the architecture.
- Estimate the cost envelope from `tokens × price/MTok × volume` and compare it to business value.
- Build the differentiator, buy the commodity, and don't force AI onto provable-exact deterministic work.
- Use the SCOPE framework; finish Criteria before choosing the Path.
- Ship the smallest first version that tests the metric, with review gates on high-stakes or irreversible steps.
