# Stop Reasons & Errors

Every stop_reason and HTTP error with cause, detection and handling code — the reliability reference for the agentic loop.

import { Tabs, TabItem, Steps } from '@prosefly/astro-components';

Reliability items hinge on two enumerations: the `stop_reason` values a loop must branch on, and the HTTP errors a client must classify into retry-or-fix. This page covers every one with cause, detection and handling.

:::danger[The reliability rule]
Branch on **`stop_reason`** to control the loop and on the **HTTP status** to decide retry-vs-fix. Never parse prose for "done", never use an iteration cap as the primary stop, and never retry a 4xx client error.
:::

## stop_reason — every value

| Value | Cause | Detection | Handling |
| --- | --- | --- | --- |
| `end_turn` | Model finished naturally | `stop_reason == "end_turn"` | Done — return the answer |
| `tool_use` | Model wants tools run | `stop_reason == "tool_use"` | Execute all `tool_use` blocks, append `tool_result`(s) in a user message, call again |
| `max_tokens` | Output hit `max_tokens` | `stop_reason == "max_tokens"` | Truncated — continue ("Continue.") or raise the cap; never treat as complete |
| `stop_sequence` | Hit a configured stop string | `stop_reason == "stop_sequence"`; check `stop_sequence` | Done — the string was the intended terminator |
| `pause_turn` | Long server-side tool turn paused | `stop_reason == "pause_turn"` | Re-send the conversation unchanged to continue |
| `refusal` | Safety systems declined | `stop_reason == "refusal"` | Explicit fallback (rephrase, human handoff, safe message); do **not** blindly retry |

```python
def handle(r, messages):
    sr = r.stop_reason
    if sr == "tool_use":
        messages.append({"role": "assistant", "content": r.content})
        messages.append({"role": "user", "content": run_tools(r.content)})
        return "continue"
    if sr == "max_tokens":
        messages.append({"role": "assistant", "content": r.content})
        messages.append({"role": "user", "content": "Continue from where you stopped."})
        return "continue"
    if sr == "pause_turn":
        return "resend"                 # re-send unchanged
    if sr == "refusal":
        return handle_refusal(r)        # explicit fallback path
    return "done"                       # end_turn / stop_sequence
```

:::caution[max_tokens is not success]
A `max_tokens` stop means the answer is cut off. Returning it as final is the silent-failure anti-pattern — especially dangerous for structured output where the JSON is now invalid.
:::

## HTTP errors — every status

| HTTP | Type | Cause | Retry? | Handling |
| --- | --- | --- | --- | --- |
| 400 | `invalid_request_error` | Malformed request (e.g., forced `tool_choice` on Fable 5.1, `budget_tokens` on non-Haiku, bad schema) | **No** | Fix the request |
| 401 | `authentication_error` | Missing/invalid API key | **No** | Fix credentials |
| 403 | `permission_error` | Key lacks entitlement (model/feature) | **No** | Check access/entitlement |
| 404 | `not_found_error` | Model/resource does not exist (e.g., retired model) | **No** | Fix the ID / migrate |
| 413 | `request_too_large` | Payload exceeds limits | **No** | Shrink input; chunk; Files API |
| 429 | `rate_limit_error` | RPM/ITPM/OTPM exceeded | **Yes** | Backoff + jitter, honour `retry-after`; throttle proactively |
| 500 | `api_error` | Server-side error | **Yes** | Backoff + jitter |
| 529 | `overloaded_error` | Capacity overloaded | **Yes** | Backoff; consider fallback to a newer-or-equal model |

```text
Retry?  ── status in {429, 500, 502, 503, 529} ──► YES: exponential backoff + jitter, honour retry-after
        └─ status in {400, 401, 403, 404, 413}  ──► NO: fix the request/credentials/entitlement
```

<Tabs>
  <TabItem label="Python">
```python
import time, random
from anthropic import APIStatusError, RateLimitError, APIConnectionError

RETRYABLE = {429, 500, 502, 503, 529}

def call(fn, **params):
    for attempt in range(6):
        try:
            return fn(**params)
        except RateLimitError as e:
            wait = float(e.response.headers.get("retry-after", 0)) or min(60, 2 ** attempt)
            time.sleep(wait + random.uniform(0, 0.5))
        except APIStatusError as e:
            if e.status_code in RETRYABLE:
                time.sleep(min(60, 2 ** attempt) + random.uniform(0, 0.5))
            else:
                raise                     # 4xx → fix, don't retry
        except APIConnectionError:
            time.sleep(min(60, 2 ** attempt) + random.uniform(0, 0.5))
    raise RuntimeError("exhausted retries")
```
  </TabItem>
  <TabItem label="TypeScript">
```typescript
import { APIError } from '@anthropic-ai/sdk';
const RETRYABLE = new Set([429, 500, 502, 503, 529]);
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));

export async function call<T>(fn: () => Promise<T>): Promise<T> {
  for (let attempt = 0; attempt < 6; attempt++) {
    try {
      return await fn();
    } catch (err) {
      if (err instanceof APIError && RETRYABLE.has(err.status ?? 0)) {
        const ra = Number(err.headers?.['retry-after']) || Math.min(60, 2 ** attempt);
        await sleep((ra + Math.random() * 0.5) * 1000);
        continue;
      }
      throw err; // 4xx → fix
    }
  }
  throw new Error('exhausted retries');
}
```
  </TabItem>
</Tabs>

## Streaming errors

An `error` event can arrive mid-stream (commonly `overloaded_error`). Handle it like the HTTP status: retryable overload/timeout → backoff and restart the stream; malformed request → fix.

```text
event: error
data: {"type":"error","error":{"type":"overloaded_error","message":"Overloaded"}}
```

## Tool execution errors (not HTTP)

A tool that fails is **not** an API error — the API call succeeded. Return a structured `tool_result` with `is_error: true`:

```json
{ "type": "tool_result", "tool_use_id": "toolu_01", "is_error": true,
  "content": "{\"category\":\"not_found\",\"retryable\":false,\"message\":\"No order ORD-999\"}" }
```

| Category | Meaning | Model should |
| --- | --- | --- |
| `not_found` | Resource absent | Tell the user, ask for a valid ID |
| `invalid_input` | Bad arguments | Correct and retry |
| `transient` | Temporary (timeout) | Retry once, then report |
| `forbidden` | Not permitted | Stop; escalate |

Never return empty success on failure (silent empty-success anti-pattern).

## Decision quick reference

| Symptom | Diagnosis | Action |
| --- | --- | --- |
| Loop never ends | Using prose/iteration cap to stop | Branch on `stop_reason` |
| JSON truncated | `max_tokens` reached | Raise cap / continue; validate |
| "It stopped and paused" | `pause_turn` | Re-send unchanged |
| Safety decline treated as crash | `refusal` mishandled | Explicit fallback |
| 429 storms | Retrying without backoff or ignoring `retry-after` | Backoff + jitter + header; raise tier |
| Retrying a 400 forever | Retrying a client error | Fix the request |
| Thinking gone after fallback | Fell back to an older model | Only migrate to newer-or-equal |
| Tool "worked" but returned nothing | Empty success | Structured `is_error` |

## Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| "Any error should be retried" | Only 429/5xx/529; fix 4xx | Retry-storm distractor |
| "`max_tokens` is a complete answer" | It is truncation | Silent-failure anti-pattern |
| "Parse the text to know it's done" | Branch on `stop_reason` | Prose-parsing anti-pattern |
| "Cap iterations to stop the loop" | Cap is a guard; `stop_reason` stops it | Iteration-cap anti-pattern |
| "`refusal` is a server error" | It is a safety decline; handle explicitly | Refusal-mishandling distractor |
| "A tool failure is an API error" | Tool errors use `is_error` on the result | Error-layer confusion |
| "529 means give up" | Backoff; consider a fallback model | Availability distractor |

## Scenario walkthrough

An agent on Opus 5 in production intermittently: (a) returns cut-off JSON, (b) hits 529s under load, (c) once returned a safety refusal shown to users as "Error 500", and (d) an engineer added `tool_choice: "any"` after switching one flow to Fable 5.1 and now gets 400s.

<Steps>
1. **Cut-off JSON** — `stop_reason: max_tokens`. Raise `max_tokens` and/or continue; validate before use.
2. **529s** — retryable; exponential backoff + jitter honouring `retry-after`; after repeated 529, fall back to a newer-or-equal model (not down).
3. **Refusal shown as 500** — `stop_reason: refusal` was mishandled; route to an explicit fallback message, not a generic error.
4. **400 on Fable 5.1** — forced `tool_choice` is unsupported (400, not retryable); switch to `auto`+instruction, `strict: true`, or `output_config.format`.
</Steps>

Rejected alternatives: retrying the 400 (client error — fix it), treating truncated JSON as valid (silent failure), and surfacing the refusal as a 500 (refusal-mishandling).

## Key takeaways

- `stop_reason` values: `end_turn`, `tool_use`, `max_tokens`, `stop_sequence`, `pause_turn`, `refusal` — branch on each.
- Retry 429/5xx/529 with backoff + jitter + `retry-after`; **fix** 400/401/403/404/413.
- `max_tokens` = truncation, `refusal` = safety decline — both need explicit handling, not silent pass-through.
- Tool failures use structured `is_error`, never empty success or an HTTP error.
- On repeated 529, fall back only to a newer-or-equal model to keep Fable 5.1 thinking blocks valid.
