# D6 · Performance, Latency and Cost

Prompt caching, Batch, Flex processing, fast mode, predicted outputs, latency budget decomposition, model downshifting and token accounting with worked cost calculations.

import { Accordions, AccordionItem } from '@prosefly/astro-components';

This domain is about **12%** of the OAI-API mock – roughly **7 of 60 items** – and mirrors the Academy *Optimize AI Application Performance* course (30 min). It tests whether you can make an AI application fast, reliable and affordable in production: caching stable prompts, batching non-urgent work, choosing the right processing tier, decomposing a latency budget, downshifting the model when quality allows, and doing the token math that decides cost.

## What you need to know

Performance work is a set of levers you pull *after* the app works and an eval protects quality. **Prompt caching** cuts cost and latency when a long prefix repeats. **Batch** trades latency for a large discount on non-urgent bulk jobs. **Flex processing** offers cheaper compute for latency-tolerant work; **fast mode** targets low latency for interactive work. **Predicted outputs** speed up responses that are largely known in advance. You decompose a **latency budget** to find the real bottleneck, **downshift** the model to the cheapest one that still passes the eval, and back every decision with **token accounting** – cost equals `tokens ÷ 1,000,000 × price/MTok`, per request and at volume.

## Learning objectives

By the end of this page you should be able to:

1. Apply **prompt caching** and structure prompts so the cache hits.
2. Choose between **Batch**, **Flex processing** and **fast mode** by latency tolerance.
3. Use **predicted outputs** where much of the output is known.
4. **Decompose a latency budget** to find the true bottleneck.
5. **Downshift** the model and reasoning effort without breaking the eval.
6. Do **token accounting** and compute cost from real prices.

---

## 6.1 The optimisation order

```text
1. Make it work           (correct output)
2. Protect it with an eval (so optimisation can't silently break quality)
3. THEN optimise:
      cost  ── caching, downshift model/effort, batch, flex
      speed ── stream, fast mode, predicted outputs, cut tokens
4. Re-run the eval after every change
```

Optimising before you have an eval means you cannot tell whether a cheaper model quietly ruined quality. Order matters.

:::tip[Assessment signal]
"Reduce cost without losing quality", "the app is too slow", "cheapest option that still meets the requirement" (MOST cost-effective) are performance items. The correct answer usually names a specific lever and preserves the eval.
:::

## 6.2 Prompt caching

When many requests share a long, identical **prefix** (a big system prompt, a fixed instruction block, a stable document), caching lets you avoid reprocessing it – cached input tokens are cheaper and faster.

```text
[  long stable prefix (cached)  ][ variable user input (not cached) ]
        pay once, reuse                pay per request
```

**Design rule:** put the stable content **first** and the variable content **last**, so the longest possible prefix is cacheable. Reordering a prompt to move variable content to the front destroys the cache hit.

## 6.3 Batch, Flex and fast mode

| Tier | Latency | Cost | Use for |
| --- | --- | --- | --- |
| **Batch** | Hours (asynchronous) | Large discount | Bulk offline jobs: nightly classification, backfills, embeddings |
| **Flex processing** | Slower than standard | Cheaper than standard | Latency-tolerant work you still want back reasonably soon |
| **Standard** | Normal | Baseline | Typical interactive requests |
| **Fast mode** | Lowest | Optimised for speed | Interactive UX where latency is the priority |

```text
Can the result wait hours?  ── yes ──► Batch (cheapest for bulk)
        │ no
Is some delay acceptable?   ── yes ──► Flex processing
        │ no
Is a user waiting live?     ── yes ──► fast mode + streaming
```

## 6.4 Predicted outputs

When you already know most of the output – editing a document, regenerating a file with small changes, reformatting known text – **predicted outputs** let you supply the expected content so the model confirms/edits it, cutting latency substantially versus regenerating from scratch.

Typical fit: "apply this small change to this large file", where 95% of the output equals the input.

## 6.5 Decomposing a latency budget

"It's slow" is not actionable. Break the wall-clock time into parts and attack the biggest.

| Component | Typical cause | Lever |
| --- | --- | --- |
| Time-to-first-token | High reasoning effort, cold cache, big prompt | Lower effort, cache prefix, trim prompt |
| Generation time | Many output tokens, big model | Fewer output tokens, downshift model, predicted outputs |
| Tool/retrieval time | Slow function or vector search | Optimise the tool, parallelise, cache |
| Network/queueing | Region, tier | Fast mode, closer region |

```text
Total latency = TTFT + generation + tool/retrieval + network
                 ▲          ▲            ▲
            biggest term? attack THAT, then re-measure.
```

## 6.6 Model and effort downshifting

The cheapest reliable win is usually **downshifting**: move from Astra/Sol to Terra/Luna, or lower reasoning effort, and confirm the eval still passes.

| From | To | When |
| --- | --- | --- |
| `gpt-6-astra` | `gpt-5.6-sol` / `terra` | Eval shows the cheaper model meets the bar |
| `gpt-5.6-sol` | `gpt-5.6-terra` | General work; Terra passes |
| `gpt-5.6-terra` | `gpt-5.6-luna` | Clear, repeatable, high-volume tasks |
| `high` effort | `medium` / `low` | Failures were not reasoning-depth failures |

Every downshift is a change → re-run the eval (D3).

## 6.7 Token accounting

Cost is arithmetic. Prices per **million tokens** (September 2026):

| Model | Input $/MTok | Output $/MTok |
| --- | --- | --- |
| `gpt-6-astra` | $10 | $50 |
| `gpt-5.6-sol` | $4 | $20 |
| `gpt-5.6-terra` | $2 | $12 |
| `gpt-5.6-luna` | $0.20 | $1.20 |

**Worked sum – choosing a model for a summariser.** 200,000 requests/day, 2,000 input + 300 output tokens each.

```text
Per day: input  200,000 × 2,000 = 400,000,000 tok = 400 MTok
         output 200,000 ×   300 =  60,000,000 tok =  60 MTok

Terra:  400 × $2  + 60 × $12  = $800 + $720   = $1,520/day  (~$45,600/mo)
Luna:   400 × $0.20 + 60 × $1.20 = $80 + $72  =   $152/day  (~$4,560/mo)
Sol:    400 × $4  + 60 × $20  = $1,600 + $1,200 = $2,800/day (~$84,000/mo)
```

If the eval shows Luna meets the summarisation bar, choosing Luna over Terra saves ~$1,368/day (~$41k/month) for the same quality. That single downshift, justified by an eval, is the highest-leverage cost decision in the domain.

## Decision framework

Use the **CLIP** framework to optimise without breaking quality.

| Letter | Step | Question |
| --- | --- | --- |
| **C** | Cache | Is there a long stable prefix to put first and cache? |
| **L** | Latency tier | Batch (hours), Flex (soon), or fast mode (now)? |
| **I** | Inference size | Can you downshift model/effort and still pass the eval? |
| **P** | Prove | Did the eval score hold after each change? |

The rule the course stresses: **change one lever, re-measure, keep only what holds quality.**

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Optimising before an eval exists | Cost pressure | Build the eval first so cuts can't hide quality loss |
| Putting variable content first in the prompt | Natural writing order | Stable prefix first so caching hits |
| Using standard/real-time for a nightly backlog | Habit | Use Batch for non-urgent bulk; large discount |
| Defaulting to the biggest model | "Safest" | Downshift to the cheapest model the eval allows |
| Guessing the latency bottleneck | "It's just slow" | Decompose the latency budget; attack the biggest term |
| Regenerating a near-identical large output | Not knowing predicted outputs | Use predicted outputs when most output is known |
| Ignoring cost until the bill arrives | Tokens feel abstract | Do `tokens × price/MTok × volume` up front |
| Cutting quality to save cost silently | No re-measurement | Re-run the eval after every optimisation |

## Scenario challenge

**Scenario.** Your product has two AI features. Feature A is an interactive chat where users wait for answers; it uses a 6,000-token system prompt plus a small user message, runs on `gpt-5.6-sol` at `high` effort, and feels sluggish. Feature B classifies a nightly backlog of 3M documents on `gpt-5.6-sol`, and the monthly bill is alarming. Leadership wants both faster and cheaper without hurting quality, and you have evals for both.

**Expert reasoning trace.**

1. **Feature A is a latency problem; decompose it.** The 6,000-token system prompt is a fixed prefix reprocessed every turn, inflating time-to-first-token. Move stable content to the front and enable **prompt caching** so the prefix is cheap and fast after the first hit. Then question `high` effort: if the eval passes at `medium`, downshift effort to cut generation time. Stream the answer so it feels instant.
2. **Feature B is a cost problem with no latency need.** A nightly backlog can wait hours, so move it to the **Batch API** for the large discount. Then test **downshifting** the model: if the eval shows `gpt-5.6-luna` classifies at target, the price drop from Sol ($4/$20) to Luna ($0.20/$1.20) is ~20× on input and ~17× on output.
3. **Token math on Feature B.** Say 3M docs × (800 input + 40 output). Sol: `2,400 MTok × $4 + 120 MTok × $20 = $9,600 + $2,400 = $12,000` per run. Luna: `2,400 × $0.20 + 120 × $1.20 = $480 + $144 = $624`. Batch adds a further discount on top. The Sol→Luna downshift alone is the dominant saving.
4. **Protect quality.** Every change – caching, effort downshift, model downshift, Batch – is validated by re-running the respective eval. If Luna misses the classification bar, stop at Terra.
5. **Do not confuse the two features.** Batch would ruin Feature A (users wait live); fast mode/caching would not help Feature B's cost. Match the lever to the constraint.

**Exam-correct decision:** for A, cache the stable prefix, downshift effort if the eval allows, and stream; for B, use Batch and downshift the model to the cheapest that passes the eval. **Not** one blanket setting for both, **not** any cut without re-running the eval.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "Use the biggest model to be safe" | Best model bias | MOST cost-effective = cheapest that still passes the eval |
| "Batch everything to save money" | Batch is cheap | Batch adds hours of latency; wrong for interactive features |
| "Caching won't help our variable prompts" | Prompts look unique | A long stable prefix (system prompt/doc) is cacheable if placed first |
| "It's slow, upgrade the model" | Bigger feels faster | Decompose latency; a bigger model often adds latency |
| "Lower cost, no need to re-check quality" | Saving feels safe | Every optimisation is a change; re-run the eval |
| "Regenerate the whole file for a small edit" | Simplest to code | Predicted outputs cut latency when most output is known |
| "Cost is roughly the same across models" | Tokens seem cheap | Luna is ~20× cheaper than Sol on input; do the math |

## Practice questions

Each item states how many responses to select. Commit before revealing.

<Accordions>
  <AccordionItem title="Q1 · Many requests share a 5,000-token system prompt followed by a short user message. What reduces cost and latency MOST directly? (Select one)">
    A. Switch to `gpt-6-astra`.
    B. Enable prompt caching with the stable system prompt placed first so the long prefix is cached.
    C. Add more output tokens.
    D. Move the user message to the front of the prompt.

    **Answer: B.** A long stable prefix placed first is the ideal prompt-caching case, cutting cost and time-to-first-token. A bigger model (A) costs more, more output (C) is slower, and moving the variable message first (D) destroys the cacheable prefix.
  </AccordionItem>

  <AccordionItem title="Q2 · A nightly job classifies 5M documents and latency does not matter. Which processing choice is MOST cost-effective? (Select one)">
    A. Real-time standard requests.
    B. The Batch API, which trades latency for a large discount on bulk work.
    C. Fast mode.
    D. Streaming.

    **Answer: B.** Non-urgent bulk work is the Batch use case: hours of latency for a large discount. Standard (A) and fast mode (C) pay for speed you do not need, and streaming (D) is a UX feature, not a cost tier.
  </AccordionItem>

  <AccordionItem title="Q3 · A summariser runs 100,000 times/day at 1,000 input and 200 output tokens. What is the daily cost on `gpt-5.6-luna`? (Select one)">
    A. About $0.44.
    B. About $4.40.
    C. About $44.
    D. About $440.

    **Answer: C.** Input: `100,000 × 1,000 = 100 MTok × $0.20 = $20.00`. Output: `100,000 × 200 = 20 MTok × $1.20 = $24.00`. Total ≈ `$44.00/day`. The discriminator is doing `MTok × price/MTok` exactly; A and B are 100× and 10× too low, and D is 10× too high.
  </AccordionItem>

  <AccordionItem title="Q4 · An interactive feature feels slow. Before changing anything, what should you do? (Select one)">
    A. Immediately upgrade to the largest model.
    B. Decompose the latency budget (time-to-first-token, generation, tool/retrieval, network) and attack the largest term.
    C. Add more output tokens.
    D. Switch to Batch.

    **Answer: B.** You cannot fix latency without knowing which component dominates, so decompose first. Upgrading blindly (A) often adds latency, more output (C) is slower, and Batch (D) is wrong for an interactive feature.
  </AccordionItem>

  <AccordionItem title="Q5 · You are editing a large document where 95% of the output equals the input. Which feature cuts latency MOST? (Select one)">
    A. Predicted outputs, supplying the expected content so the model confirms/edits rather than regenerating.
    B. Higher reasoning effort.
    C. A larger context window.
    D. Batch processing.

    **Answer: A.** Predicted outputs speed responses whose output is mostly known in advance, ideal for small edits to large files. Higher effort (B) is slower, context size (C) does not address regeneration cost, and Batch (D) adds latency.
  </AccordionItem>

  <AccordionItem title="Q6 · An eval shows `gpt-5.6-luna` classifies at the target that `gpt-5.6-sol` currently meets. What is the MOST cost-effective action? (Select two)">
    A. Downshift from Sol to Luna for this task.
    B. Re-run the eval after the switch to confirm quality holds.
    C. Keep Sol because bigger is safer.
    D. Upgrade to `gpt-6-astra`.
    E. Disable the eval to save compute.

    **Answer: A and B.** If Luna passes the eval it is far cheaper, so downshift and re-verify. Keeping Sol (C) or upgrading to Astra (D) wastes money for no measured gain, and disabling the eval (E) removes the quality guard.
  </AccordionItem>

  <AccordionItem title="Q7 · Which work is the WRONG fit for the Batch API? (Select one)">
    A. Nightly re-embedding of a document corpus.
    B. A live chat where the user waits for each reply.
    C. A weekend backfill of classifications.
    D. Bulk generation of product descriptions overnight.

    **Answer: B.** Batch introduces hours of latency, which is unacceptable when a user waits live. Nightly embeddings (A), a weekend backfill (C) and overnight bulk generation (D) are all latency-tolerant and fit Batch well.
  </AccordionItem>

  <AccordionItem title="Q8 · A team lowers cost by switching to a cheaper model but does not re-check quality. What is the risk? (Select one)">
    A. None; cheaper is always fine.
    B. The cheaper model may silently miss the quality bar; without re-running the eval the regression ships unnoticed.
    C. The API will reject the request.
    D. Latency will always increase.

    **Answer: B.** A downshift is a change that can degrade quality; only re-running the eval catches it. Cheaper is not always fine (A), the API does not reject a valid model (C), and a cheaper model often lowers, not raises, latency (D).
  </AccordionItem>

  <AccordionItem title="Q9 · What is the correct order of operations when a working feature is too expensive? (Select one)">
    A. Optimise first, then build an eval if there is time.
    B. Ensure an eval protects quality, then apply cost levers (caching, downshift, batch, flex), re-measuring after each.
    C. Immediately switch to the cheapest model and ship.
    D. Remove the feature.

    **Answer: B.** Optimisation must be guarded by an eval so cuts do not silently break quality; then apply levers and re-measure. Optimising before the eval (A) or switching blindly (C) risks unnoticed regressions, and removing the feature (D) is not optimisation.
  </AccordionItem>

  <AccordionItem title="Q10 · A latency-tolerant workload wants lower cost than standard but cannot wait the hours Batch requires. Which tier fits? (Select one)">
    A. Fast mode.
    B. Flex processing, which is cheaper than standard for latency-tolerant work returned reasonably soon.
    C. Batch.
    D. Streaming.

    **Answer: B.** Flex processing sits between standard and Batch: cheaper than standard, faster than Batch, for latency-tolerant work. Fast mode (A) optimises for speed at higher cost, Batch (C) is too slow here, and streaming (D) is a UX feature.
  </AccordionItem>

  <AccordionItem title="Q11 · Time-to-first-token is the dominant latency term for an interactive feature on `high` effort. Which TWO levers directly help? (Select two)">
    A. Lower reasoning effort if the eval still passes.
    B. Cache the long stable prompt prefix so it is not reprocessed each turn.
    C. Increase max output tokens.
    D. Switch to a bigger model.
    E. Add more retrieved chunks.

    **Answer: A and B.** High effort and reprocessing a big prefix both inflate time-to-first-token, so lowering effort and caching the prefix attack it directly. More output tokens (C) affect generation time, a bigger model (D) usually adds latency, and more chunks (E) add tokens.
  </AccordionItem>

  <AccordionItem title="Q12 · A workload of 500,000 requests/day uses 1,500 input and 500 output tokens on `gpt-5.6-terra`. Roughly what is the daily cost? (Select one)">
    A. About $180.
    B. About $4,500.
    C. About $45,000.
    D. About $450,000.

    **Answer: B.** Input: `500,000 × 1,500 = 750 MTok × $2 = $1,500`. Output: `500,000 × 500 = 250 MTok × $12 = $3,000`. Total ≈ `$4,500/day`. Option A drops a factor, and C and D are 10× and 100× too high.
  </AccordionItem>
</Accordions>

## Key takeaways

- Optimise only after an eval protects quality, and re-run it after every change.
- Prompt caching cuts cost and latency when a long stable prefix is placed first.
- Match the tier to latency tolerance: Batch (hours), Flex (soon), fast mode (now); stream for waiting users.
- Predicted outputs speed responses whose output is largely known, such as small edits to large files.
- Decompose the latency budget and attack the largest term rather than guessing.
- Downshift model and reasoning effort to the cheapest option the eval allows.
- Cost is arithmetic: `tokens ÷ 1,000,000 × price/MTok × volume` – a justified model downshift is usually the biggest saving.
