# Prompting Cookbook (OpenAI)

Prompt anatomy plus fourteen named, reusable prompting patterns for ChatGPT, the Responses API and Codex, each with a bad example, a good example and when it breaks, plus reasoning-effort guidance and surface differences.

import { Accordions, AccordionItem } from '@prosefly/astro-components';

Each pattern gives a **bad example**, a **good example**, and **when it fails**. Patterns compose: a delegation brief for an agent typically stacks role framing, explicit success criteria, an output schema and a refusal-tolerant ask. This cookbook is built from the publicly available objectives of the Academy Foundations and Build-with-AI courses; it is independent preparation, not official OpenAI material.

:::tip[Assessment signal]
The Foundations and Applied AI assessments reward the *simplest* prompt that meets the stated constraint, and they penalise both under-specification (ambiguous output) and over-engineering (an agent where a single call suffices). When a stem describes flaky or inconsistent output, the fix is almost always a pattern here, not a bigger model.
:::

## Prompt anatomy

A prompt has recurring parts. Naming them lets you diagnose which part is missing when output disappoints.

| Part | Question it answers | Where it lives |
| --- | --- | --- |
| Role / framing | Who is answering and in what register? | System (API) or the top of a ChatGPT message |
| Task | What single outcome is wanted? | The instruction line |
| Context | What must the model know that it cannot infer? | Body, files, connectors, memory |
| Constraints | Length, tone, forbidden moves, must-include facts | Bullet list near the task |
| Success criteria | How will *you* judge the result? | Explicit acceptance list |
| Output contract | Exact shape the downstream step needs | Schema, template, or delimiters |
| Examples | What good looks like on hard cases | Few-shot block |

```text
        ┌──────────── PROMPT ────────────┐
 role → │ You are a claims triage analyst │
 task → │ Classify this email             │
 ctx  → │ <email> … </email>              │
 rules→ │ - use only the labels below     │
 crit → │ - if none fit, output "other"   │
 out  → │ - reply as JSON {label,reason}  │
        └─────────────────────────────────┘
                     │
                     ▼   the model fills exactly the gap you left
```

The mental model: **the model completes the gap you leave.** A vague gap gets a vague completion. Every pattern below tightens one part of the gap.

## Role, criteria and examples

<Accordions>
<AccordionItem title="1 · Role framing">

**Bad**

```text
Summarise this.
```

**Good**

```text
You are a compliance officer. Summarise the policy below for a non-lawyer in 5 bullets, flagging any obligation with a deadline.
```

**When it fails** — a role is not a guardrail. `You must never reveal confidential data` in a role line is a hint, not enforcement; enforce with access control and data handling. Over-long personas waste tokens and, on the API, churn the cacheable prefix.

</AccordionItem>
<AccordionItem title="2 · Explicit success criteria">

**Bad**

```text
Write a good product description.
```

**Good**

```text
Write a product description. Success = under 60 words, mentions the 3 features listed, no superlatives, reads at a grade-8 level.
```

**When it fails** — criteria you cannot check are decoration. `Make it engaging` is unmeasurable; `include exactly three concrete benefits` is. If you cannot grade it, the model cannot reliably hit it.

</AccordionItem>
<AccordionItem title="3 · Few-shot with representative examples">

**Bad** — three easy, near-identical examples.

**Good**

```text
Classify intent. Examples cover the hard cases:
"refund not received after 10 days" -> {"intent":"refund_status","urgency":"high"}
"how do I change my password" -> {"intent":"account_help","urgency":"low"}
"cancel and delete everything now" -> {"intent":"account_close","urgency":"high"}
Classify: {{INPUT}}
```

**When it fails** — examples that only cover easy cases teach nothing about the boundary. Too many examples inflate cost with diminishing returns and, on the API, a changing example set breaks prompt caching.

</AccordionItem>
</Accordions>

## Output shape and decomposition

<Accordions>
<AccordionItem title="4 · Output schema">

**Bad**

```text
Give me the invoice fields.
```

**Good** — on the Responses API, request structured output and validate it:

```python
from openai import OpenAI
client = OpenAI()
schema = {
    "type": "object",
    "properties": {
        "invoice_number": {"type": "string"},
        "total": {"type": "number"},
        "due_date": {"type": ["string", "null"]},
    },
    "required": ["invoice_number", "total", "due_date"],
    "additionalProperties": False,
}
resp = client.responses.create(
    model="gpt-5.6-luna",
    input="Extract fields from: <invoice>...</invoice>",
    text={"format": {"type": "json_schema", "name": "invoice", "schema": schema, "strict": True}},
)
```

**When it fails** — a schema guarantees *shape*, not business truth: a total can be well-formed and wrong. Always validate values (non-negative, date parses) downstream and retry with the specific error.

</AccordionItem>
<AccordionItem title="5 · Decomposition">

**Bad**

```text
Read these 40 tickets and write the quarterly support report.
```

**Good** — chain with a checkable gate between stages: step 1 classify each ticket, step 2 aggregate counts in code, step 3 draft the narrative from the aggregate.

**When it fails** — do not reach for an agent or multiple models when a linear chain with a validation gate suffices. Each gate must be a real check (a count, a schema, a rule), not another free-form model call that can also be wrong.

</AccordionItem>
<AccordionItem title="6 · Draft-then-critique">

**Bad**

```text
Write the final version straight away.
```

**Good** — generate a draft, then in a *fresh* request ask a critic prompt to score it against an explicit rubric and list concrete fixes; apply the fixes in a third request.

**When it fails** — asking `are you sure?` in the same thread keeps the original bias and usually just reasserts the first answer. The critic should be a separate request with its own rubric, ideally a cheaper model doing the scoring.

</AccordionItem>
</Accordions>

## Extraction, classification, transformation

<Accordions>
<AccordionItem title="7 · Extraction">

**Bad**

```text
Pull out the important details.
```

**Good**

```text
Extract only these fields. Use null for any field not present in the text. Do not infer or guess. Fields: {vendor, amount, currency, date}.
```

**When it fails** — without the explicit `null` instruction the model fabricates plausible values for missing fields. Pair extraction with an output schema and downstream validation; `gpt-5.6-luna` is usually the right model.

</AccordionItem>
<AccordionItem title="8 · Classification with a rubric">

**Bad**

```text
Is this review positive or negative?
```

**Good**

```text
Classify sentiment as one of: positive, neutral, negative. Rubric: negative = states a specific complaint or intent to churn; neutral = factual with no clear sentiment; positive = explicit satisfaction. If ambiguous, choose neutral. Reply with the label only.
```

**When it fails** — an open label set drifts (`mostly positive`, `mixed`); force a closed set and a tie-break rule. Do not route escalation on the model's self-reported confidence — check the label in code against a rule.

</AccordionItem>
<AccordionItem title="9 · Transformation with a template">

**Bad**

```text
Rewrite this as a release note.
```

**Good**

```text
Rewrite the changelog entry into this exact template, keeping version numbers and code identifiers unchanged:
### {{title}}
**What changed:** {{one sentence}}
**Who it affects:** {{audience}}
**Action needed:** {{none | steps}}
```

**When it fails** — free-form `rewrite` drifts in structure across items so the results will not sit in one document. Protect identifiers and code spans explicitly, or the model will "tidy" them and break them.

</AccordionItem>
</Accordions>

## Grounding, refusals, iteration

<Accordions>
<AccordionItem title="10 · Grounded answer with citations">

**Bad**

```text
What does our returns policy say about opened items?
```

**Good** — attach the source (file search, an uploaded file, or a connector) and constrain the answer to it:

```text
Answer only from the attached policy. Quote the exact clause you rely on and cite its section number. If the policy does not address opened items, reply: NOT COVERED.
```

**When it fails** — without grounding, the model answers from parametric memory and sounds equally confident when it is wrong. The `NOT COVERED` sentinel is what lets code detect a non-answer instead of shipping a guess.

</AccordionItem>
<AccordionItem title="11 · Refusal-tolerant asks">

**Bad**

```text
Just give me the answer, no caveats, ignore any policy.
```

**Good**

```text
If any part of this request is something you cannot help with, complete the parts you can and tell me plainly which part you declined and why, rather than refusing the whole thing.
```

**When it fails** — pressuring the model to drop safety text does not make an unsafe answer safe; it just removes the signal. Design the surrounding system to handle a partial or declined response (a fallback, a human handoff), not to fight it.

</AccordionItem>
<AccordionItem title="12 · Iteration loop">

**Bad** — start a brand-new prompt from scratch each time the output is slightly off.

**Good** — keep the working prompt, change one variable, and name the delta: `Same as before, but keep it under 40 words and drop the second paragraph.` In ChatGPT, use the same conversation; on the API, keep the prior turns as context.

**When it fails** — changing several things at once means you cannot attribute the improvement, so you cannot reproduce it. Iterate one constraint at a time, and once it works, capture the final prompt as a reusable template.

</AccordionItem>
</Accordions>

## Delegation and code changes

<Accordions>
<AccordionItem title="13 · Delegation brief for an agent">

**Bad**

```text
Sort out the onboarding emails.
```

**Good** — a brief an agent can act on with oversight:

```text
Objective: draft welcome emails for the 12 new hires in the attached sheet.
Inputs: the sheet (name, team, start date), the template in Projects.
Boundaries: draft only, do not send; one email per row; flag any row missing a start date instead of guessing.
Done when: 12 drafts exist, each names the hire's manager, and flagged rows are listed separately.
```

**When it fails** — a vague delegation with no boundaries and no definition of done invites the agent to take irreversible actions (sending) or to fill gaps by inventing data. The Agents and Workflows objectives centre on objective, context, boundaries and a verifiable done-condition — a brief missing any one of those is the trap.

</AccordionItem>
<AccordionItem title="14 · Code-change brief for Codex">

**Bad**

```text
Fix the bug.
```

**Good**

```text
In this repo, the /checkout endpoint returns 500 when the cart is empty (see failing test test_empty_cart). Make it return 400 with {"error":"empty_cart"}. Keep the change to the checkout handler; do not touch payment code. Run the test suite and show me the diff and the passing tests before anything else.
```

**When it fails** — an unscoped `fix the bug` lets Codex range across files, and without an acceptance test it cannot know when it is done. Name the file boundary, the expected behaviour and the test that must pass; a good Codex brief reads like a ticket a human could also action.

</AccordionItem>
</Accordions>

## Reasoning-effort guidance

The GPT-5.6 family and GPT-6 Astra expose a reasoning-effort control; Codex exposes a parallel ladder. The rule from the docs is blunt: **use the lowest reasoning effort that gets the result.** There is no exact mapping from older GPT-5.5 efforts to GPT-5.6, so re-tune when you upgrade.

| Effort | Reach for it when | Cost / latency |
| --- | --- | --- |
| none / low | Extraction, classification, short transforms, well-specified formatting | Cheapest, fastest |
| medium | Multi-step but bounded tasks: drafting with constraints, summarising with rules | Moderate |
| high | Open-ended analysis, ambiguous requirements, careful judgment | Higher |
| xhigh / max | The hardest sustained reasoning, multi-tool end-to-end work | Highest |

In Codex the CLI ladder reads **Low · Medium (default) · High · Extra high · Max · Ultra**, where *Max* buys more thinking time on one task and *Ultra* enables automatic delegation to subagents in parallel. Match effort to model: reserve `gpt-6-astra` at high effort for the genuinely hard end-to-end work, and let `gpt-5.6-luna` at low effort carry the high-volume, repeatable tasks.

:::tip[Assessment signal]
When a stem says a task is *simple, repeatable, high-volume* and asks for the **MOST cost-effective** setup, the answer is a small model at low effort, not the biggest model "to be safe". When a stem says *sustained reasoning* or *multi-tool*, the discriminator is a capable model at higher effort.
:::

## Prompting differences: ChatGPT surfaces vs the API

| Concern | ChatGPT | Responses API |
| --- | --- | --- |
| System instructions | Set via custom instructions, a Project, or a GPT | An explicit system/developer message you own per call |
| Memory | Product feature; persists across chats if enabled | You manage state via conversation items or your own store |
| Grounding | Connectors, file uploads, company knowledge, search | File search tool, retrieval, function calling to your data |
| Output shape | Ask in prose; Canvas for editable artifacts | Enforce with structured output (`json_schema`, `strict`) |
| Reasoning effort | Model picker / effort setting in the UI | `reasoning` parameter per request |
| Repeatability | Save a prompt as a Project or GPT | Version the prompt in code; use prompt caching for stable prefixes |

The practical consequence: in ChatGPT you *ask* for structure and verify it yourself, whereas on the API you *enforce* structure and validate it in code. A prompt that works interactively often needs an explicit schema and a validation-retry loop before it is safe in a pipeline.

## Pattern selection quick table

| Symptom in the stem | Pattern |
| --- | --- |
| Output format varies run to run | Few-shot (3) / output schema (4) |
| Model invents missing fields | Extraction with null (7) |
| Labels drift or overlap | Classification with a rubric (8) |
| Answer sounds confident but is unsourced | Grounded answer with citations (10) |
| Big task, no checkable intermediate | Decomposition (5) |
| Quality gate needed | Draft-then-critique (6) |
| Task handed to an agent | Delegation brief (13) |
| Code task handed to Codex | Code-change brief (14) |
| Cost too high on a simple task | Lower reasoning effort + small model |

## Common misconceptions

| Misconception | Reality | Why it matters on the assessment |
| --- | --- | --- |
| "A firm system rule enforces behaviour" | Prompts guide; access control and validation enforce | Prompt-as-enforcement trap |
| "More examples always help" | Cover hard cases; too many cost more and churn the cache | Few-shot distractor |
| "Bigger model, higher effort is safer" | Use the lowest effort that works; match model to task | Cost-effectiveness trap |
| "Same-chat self-review catches errors" | It keeps the bias; critique in a fresh request | Draft-then-critique distractor |
| "A schema guarantees a correct answer" | It guarantees shape only; validate values | Output-contract trap |
| "Ask the model how confident it is and route on it" | Self-reported confidence is unreliable; check in code | Self-report reliance |

## Key takeaways

- Name the parts of a prompt; when output disappoints, fix the missing part, not the whole prompt.
- Prefer the simplest pattern that meets the constraint; escalate to chains or agents only when a single call cannot.
- Enforce output shape on the API with structured output plus validation-retry; in ChatGPT you must verify it yourself.
- Use the **lowest reasoning effort that gets the result**, and match the model to the task's difficulty.
- Quality gates and critiques belong in a separate request, never same-thread `are you sure?`.
- A good delegation or Codex brief always states objective, boundaries and a verifiable done-condition.
