# D4 · Boundaries and Guardrails

Setting the limits an agent operates inside — irreversible actions, approval gates, spend and time limits, data boundaries — and the difference between a boundary the agent respects and one the system enforces.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

Worth **16%** — about **8 of 50 items**. This domain tests whether you can bound an agent so that when it misunderstands the task (and eventually it will), the damage is contained. The single most important idea here is the distinction between a boundary the agent is *asked* to respect and one the *system* enforces. On an irreversible action, a request is not a control; enforcement is. Get that distinction right and most of the domain follows.

## What you need to know

Boundaries are the limits you place around an agent's autonomy. The highest-stakes boundaries protect against **irreversible actions** — sending, publishing, paying, deleting, changing a system of record. The primary control for these is an **approval gate**: the agent pauses and a human authorises before the action fires. Other boundaries cap **spend** and **time** so a runaway or looping agent can't burn budget or run forever, and enforce **data boundaries** so it can't move sensitive data where it shouldn't. The crucial distinction is between a **respected** boundary (the agent is instructed not to cross it and usually won't) and an **enforced** boundary (the system makes crossing it impossible). For anything irreversible or sensitive, you need enforcement — a literal, autonomous agent that misreads the goal can cross a merely-requested line without knowing it did anything wrong.

## Learning objectives

By the end of this page you should be able to:

1. **Identify** irreversible actions and place approval gates before them.
2. **Distinguish** a boundary the agent respects from one the system enforces, and know when only enforcement will do.
3. **Set** spend and time limits appropriate to a task's stakes.
4. **Define** data boundaries that stop sensitive data crossing lines it shouldn't.
5. **Design** an approval gate that pauses at the right point with enough context to decide.
6. **Choose** enforcement over instruction whenever an action is irreversible or a boundary protects against real harm.

---

## 4.1 Irreversible actions are the whole game

Every boundary decision starts by asking: *what can this agent do that I can't take back?* Reversible actions are forgiving — a bad draft is deleted, a wrong analysis re-run. Irreversible actions are not, and they are exactly where an agent's autonomy is most dangerous.

| Action | Reversible? | Boundary needed |
| --- | --- | --- |
| Draft, summarise, analyse in the workspace | Yes | Light review after |
| Save/overwrite a file | Sometimes (overwrite may lose the original) | Gate overwrites of originals |
| Send an email, post a message | No — the recipient has it | Approval gate before send |
| Publish externally | No — it's public | Approval gate + review |
| Pay, transfer, purchase | No — money moved | Enforced gate + sign-off |
| Delete, or change a system of record | No / costly | Enforced gate; often block entirely |

```text
        reversibility ◄──────────────────────────────►
        fully reversible                     irreversible
        │                                              │
        draft / analyse        overwrite        send / publish / pay / delete
        light review        gate overwrites    APPROVAL GATE (enforced)
```

:::tip[Assessment signal]
Any stem containing "send", "publish", "pay", "delete", "cannot be undone" or "external" is pointing at an irreversible action, and the correct answer will place an **approval gate** before it — never "trust the agent to be careful".
:::

## 4.2 Respected boundaries versus enforced boundaries

This is the distinction the domain is built around, and it is the most common trap.

A **respected** boundary lives in the brief: "do not email anyone outside the company", "don't spend more than an hour on this". The agent reads it and, most of the time, honours it. But it is instruction, not control — a misread objective or an unexpected path can lead the agent across the line without malice, and you find out afterward.

An **enforced** boundary lives in the *system*: the agent literally cannot send without your click; the connector has no external-send permission; a spend cap stops the run at a hard number. Crossing it is impossible, not merely discouraged.

| | Respected boundary | Enforced boundary |
| --- | --- | --- |
| Where it lives | The brief / instructions | The system / configuration |
| What it does | Asks the agent not to cross | Makes crossing impossible |
| Fails when | The agent misreads or takes an odd path | It doesn't — that's the point |
| Use for | Preferences, style, low-stakes scope | Irreversible actions, sensitive data, spend |

```text
   "please don't send externally"        vs      send tool disabled / gated
   ─────────────────────────────                 ──────────────────────────
   agent usually complies                        agent cannot send, full stop
   depends on correct understanding              independent of understanding
   OK for low-stakes preferences                 REQUIRED for irreversible / sensitive
```

**The rule:** if crossing the boundary would cause real, hard-to-reverse harm, it must be enforced, not merely requested. A well-behaved agent that respects a requested boundary 99% of the time is still an unacceptable control on an action that can't be undone.

:::tip[Assessment signal]
When a stem offers "tell the agent not to…" versus "configure the system so it can't…", and the action is irreversible or sensitive, the enforced option is correct. "Instruct it clearly" is the tempting-but-wrong answer for high-stakes boundaries.
:::

## 4.3 Approval gates: pausing at the right point

An approval gate is the enforced pause before an irreversible action. A good gate does three things: it pauses **before** the action (not after), it surfaces **enough context** to decide, and it makes the decision **cheap** for the human.

| Bad gate | Good gate |
| --- | --- |
| Pauses after the email is sent | Pauses before send, showing the full draft and recipients |
| Asks "proceed?" with no detail | Shows exactly what will happen and to whom |
| Gates everything, so humans rubber-stamp | Gates the irreversible/high-stakes steps only |
| Gates nothing, trusts the agent | Gates every irreversible action |

The design tension is real: gate too much and humans stop reading and click "approve" reflexively (gate fatigue); gate too little and something irreversible slips through. Resolve it by gating on **reversibility and stakes**, not on every step — the "one before many" checkpoint from D2 is a form of this.

## 4.4 Spend and time limits

Autonomy plus a loop equals a runaway. An agent that misunderstands "keep improving it" can iterate indefinitely, and one with paid tools can accumulate cost. Spend and time limits are enforced boundaries against this.

| Limit | Protects against | Set it by |
| --- | --- | --- |
| **Time / step limit** | Endless loops, drift on long runs | The task's realistic duration, plus margin |
| **Spend cap** | Runaway paid-tool or model cost | The value of the task; stop well below "surprising" |
| **Scope cap** (items processed) | One→many mistakes compounding | The batch you've verified the pattern on |

These are enforced, not requested: "try not to spend too much" is a respected boundary and will not stop a looping agent. A hard cap will. The point is not to be stingy but to make the *worst case* bounded and known before you start the run.

## 4.5 Data boundaries

Data boundaries stop an agent moving sensitive information where it shouldn't go — pasting confidential data into an external tool, including PII in an outbound message, or copying regulated data across a line. They connect directly to D3's reach: the tightest data boundary is *not granting reach* to the sensitive data in the first place.

| Data-boundary control | Respected or enforced | Notes |
| --- | --- | --- |
| "Don't include account numbers in emails" | Respected | Instruction; can be missed |
| Redaction/DLP that blocks the send | Enforced | Catches what instruction misses |
| Not connecting the agent to the sensitive store | Enforced | Strongest — it can't move what it can't reach |
| Scoping a connector to non-sensitive data | Enforced | Least-privilege applied to data |

The strongest data boundary is architectural: an agent cannot leak data it was never able to touch. Where reach is unavoidable, enforce the outbound control rather than trusting the instruction.

## 4.6 Matching the boundary to the stakes

Not every task needs heavy boundaries. Over-gating a low-stakes drafting task wastes attention; under-gating an irreversible one invites harm. Match the control to the reversibility and sensitivity.

```text
                         stakes / irreversibility
        low ──────────────────────────────────────► high
        internal draft      shared doc         external send / pay / delete
        │                   │                  │
        light review        spot-check +        ENFORCED approval gate +
        no gate needed      gate overwrites     spend/time caps + sign-off
```

The judgment mirrors the whole track: reversible and low-stakes → keep it light; irreversible or sensitive → enforce. The mistake in both directions is treating all tasks the same.

## Decision framework

Use the **GATE** test on every action an agent can take: **G**auge reversibility, **A**uthorise irreversible steps, **T**hreshold spend and time, **E**nforce, don't just ask.

| Step | Question | Action |
| --- | --- | --- |
| **G — Gauge** | Can this action be undone? | If no, it needs a gate — full stop |
| **A — Authorise** | Who approves before it fires? | Insert a human approval checkpoint before the irreversible step |
| **T — Threshold** | What's the worst-case cost/time? | Set an enforced spend cap and time/step limit |
| **E — Enforce** | Is this a request or a control? | For irreversible/sensitive actions, make it system-enforced, not brief-requested |

Apply it tomorrow: list what the agent *can* do, run each through GATE, and turn every irreversible or sensitive action from a respected boundary into an enforced one.

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Trusting the agent to "be careful" with irreversible actions | It usually behaves | Enforce a gate; usual behaviour isn't a control |
| Writing the boundary only in the brief | Instruction feels sufficient | Enforce high-stakes boundaries in the system |
| Gating after the action instead of before | The pause was added late | Gate before the irreversible step fires |
| Gating every step | Wanting maximum safety | Gate by reversibility/stakes; over-gating causes rubber-stamping |
| No spend or time cap on a looping-capable task | "It won't loop" | Set enforced caps sized to the task's value |
| "Don't paste sensitive data" as the only data control | Instruction is easy to write | Enforce with redaction, or don't grant reach at all |
| Approval prompt with no context | Added as an afterthought | Show what will happen and to whom, so the human can decide |
| Same boundaries for every task | Simplicity | Match the boundary to the stakes and reversibility |

## Scenario challenge

**Scenario.** Ravi runs partnerships on ChatGPT Work. He delegates an agent to "review the inbound partnership requests in our shared inbox, draft polite responses, and send the clear yes/no cases so I only handle the maybe pile". To speed things up, he adds to the brief: "Only send to companies we already have a relationship with; never send to anyone new without checking." The agent processes forty requests overnight. In the morning, Ravi finds it emailed twelve external companies — including three it had never dealt with — because it interpreted "clear yes" broadly. One message quoted internal deal terms.

**Expert reasoning trace.**

1. **Name the irreversible action.** Sending external emails is irreversible — once sent, they can't be recalled. That alone means the send should never have been a merely-respected boundary.
2. **See the respected-vs-enforced failure.** "Never send to anyone new without checking" is an instruction in the brief — a *respected* boundary. The agent misread "clear yes" and crossed it, exactly the failure mode of a respected boundary on an irreversible action. It didn't disobey maliciously; it interpreted the goal and acted.
3. **The correct control is enforcement.** The send should have been behind an *enforced* approval gate: the agent drafts, pauses, and Ravi authorises each send (or at least each send to a new company) before it fires. With enforcement, the agent's broad interpretation of "clear yes" produces drafts awaiting approval, not sent emails.
4. **Add the data boundary.** The message quoting internal deal terms is a data-boundary failure. The agent had reach to internal deal information and no enforced control stopped it leaving in an outbound message. Fix: don't grant reach to internal deal terms for this task (D3), and/or enforce an outbound check — not merely instruct "don't include internal terms".
5. **Right-size the gating.** Gating *every* draft would create rubber-stamp fatigue across forty items. Gate by stakes: auto-draft everything, enforce approval on all external sends (the irreversible step), and require explicit sign-off for any new company (the highest-stakes subset). Time/scope caps bound the overnight run so it can't process an unbounded pile.
6. **Reframe the lesson.** Ravi treated a control problem as a wording problem. No amount of clearer instruction makes a respected boundary safe for an irreversible action — the fix is to move the boundary from the brief into the system.

**The point.** The harm came from relying on a *respected* boundary where an *enforced* one was required. Sending is irreversible; internal deal terms are sensitive; both needed system-level controls (approval gate on send, no reach to internal terms), not a more strongly-worded brief.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "Just instruct the agent clearly not to…" | Clear instructions feel like control | On irreversible/sensitive actions, only *enforcement* is a control |
| "It behaved fine in testing, so trust it" | Usual behaviour reassures | A misread on the wrong run is unrecoverable; enforce the gate |
| "Gate every action to be safe" | Maximum caution seems safest | Over-gating causes rubber-stamping; gate by reversibility/stakes |
| "Add 'don't overspend' to the brief" | Easy to write | A respected cap won't stop a loop; set an enforced spend limit |
| "The draft mentioned internal terms — reword the brief" | Looks like a wording fix | It's a data-boundary/reach failure; enforce or remove reach |
| "Approve after sending to save a step" | Feels efficient | A gate after an irreversible action controls nothing |

## Practice questions

Each item states how many responses to select. Attempt before revealing.

<Accordions>
  <AccordionItem title="Q1 · Which action MOST requires an enforced approval gate before it fires? (Select one)">
    A. Drafting a summary in the workspace.
    B. Sending an email to an external customer.
    C. Analysing a spreadsheet you supplied.
    D. Rewriting a paragraph you'll read next.

    **Answer: B.** Sending an external email is irreversible, so it needs an enforced gate before it fires. Drafting (A), analysing (C) and rewriting (D) are reversible and safe with light review — nothing has left the workspace.
  </AccordionItem>

  <AccordionItem title="Q2 · What is the key difference between a respected boundary and an enforced boundary? (Select one)">
    A. Respected boundaries are written in a nicer tone.
    B. A respected boundary asks the agent not to cross a line; an enforced boundary makes crossing it impossible.
    C. Enforced boundaries only apply to large models.
    D. There is no real difference.

    **Answer: B.** A respected boundary is instruction the agent usually honours; an enforced boundary is a system control that removes the possibility of crossing. Tone (A) is irrelevant, enforcement isn't model-specific (C), and the difference is central, not cosmetic (D).
  </AccordionItem>

  <AccordionItem title="Q3 · For an irreversible action, why is 'instruct the agent clearly not to do it' insufficient? (Select one)">
    A. It isn't insufficient; clear instructions always work.
    B. A misread objective or unexpected path can lead the agent across a merely-requested line, and the action can't be undone.
    C. Instructions cost more tokens.
    D. Agents ignore all instructions.

    **Answer: B.** Instruction depends on correct understanding, and an autonomous agent can cross a requested line by misreading the goal — unacceptable when the action is irreversible. Clear instructions don't always hold (A). Token cost (C) is irrelevant, and agents don't ignore all instructions (D) — they can simply misinterpret them.
  </AccordionItem>

  <AccordionItem title="Q4 · A good approval gate should… (Select one)">
    A. Pause after the action, to avoid slowing the agent.
    B. Pause before the irreversible action and show what will happen and to whom.
    C. Ask 'proceed?' with no detail.
    D. Gate every single step equally.

    **Answer: B.** A gate must pause before the irreversible step and surface enough context — the draft and recipients — for a real decision. Pausing after (A) controls nothing. A detail-free prompt (C) invites blind approval. Gating every step (D) causes rubber-stamp fatigue.
  </AccordionItem>

  <AccordionItem title="Q5 · An agent capable of iterating and using paid tools is told 'don't spend too much'. What is the problem? (Select one)">
    A. Nothing; the instruction is enough.
    B. 'Don't spend too much' is a respected boundary that won't stop a looping run; set an enforced spend cap.
    C. Paid tools can't loop.
    D. Spend limits slow the model down.

    **Answer: B.** A vague requested limit won't halt a runaway; an enforced spend cap makes the worst case bounded and known. The instruction alone is not enough (A). Loops can occur with paid tools (C), and an enforced cap doesn't slow the model — it stops a runaway (D).
  </AccordionItem>

  <AccordionItem title="Q6 · What is the STRONGEST data boundary against an agent leaking a sensitive internal document? (Select one)">
    A. Instructing it not to share the document.
    B. Not granting the agent reach to the sensitive document at all.
    C. Asking it to summarise the document carefully.
    D. Using a larger model.

    **Answer: B.** The strongest data boundary is architectural: an agent cannot leak what it was never able to reach. Instruction (A) can be missed. Careful summarisation (C) still touches the data. Model size (D) doesn't create a data boundary.
  </AccordionItem>

  <AccordionItem title="Q7 · Why can gating every action be as harmful as gating none? (Select one)">
    A. It isn't; more gates are always better.
    B. Excessive gates cause humans to approve reflexively, so a real risk slips through unnoticed.
    C. Gates disable the agent's tools.
    D. Gates change the model.

    **Answer: B.** Over-gating produces rubber-stamp fatigue, defeating the purpose of the gate when it matters. More gates aren't always better (A). Gates don't disable tools (C) or change the model (D).
  </AccordionItem>

  <AccordionItem title="Q8 · Which TWO boundaries should be system-enforced rather than merely written in the brief? (Select two)">
    A. A preference for bullet points over prose.
    B. An approval gate before any external send.
    C. A hard spend cap on a task that uses paid tools.
    D. A suggestion to keep the tone friendly.
    E. A preference to finish before lunch.

    **Answer: B and C.** An approval gate on irreversible sends (B) and a hard spend cap on paid-tool use (C) protect against real, hard-to-reverse harm and must be enforced. Style (A), tone (D) and a soft timing wish (E) are preferences — respected boundaries are fine.
  </AccordionItem>

  <AccordionItem title="Q9 · An agent overwrites an original file while 'improving' it, losing the prior version. Which boundary would have prevented this? (Select one)">
    A. A friendlier tone.
    B. An enforced gate on overwriting originals, or saving to a new version instead.
    C. A larger context window.
    D. A web-search tool.

    **Answer: B.** Overwriting can be irreversible, so gating overwrites or forcing versioned saves preserves the original. Tone (A), context window (C) and web search (D) do nothing to protect the file.
  </AccordionItem>

  <AccordionItem title="Q10 · A task is a low-stakes internal draft you'll read before doing anything with it. What boundary posture fits? (Select one)">
    A. Heavy enforced gates on every step.
    B. Light review after the run; no gate needed because nothing irreversible happens.
    C. A spend cap of zero so it can't run.
    D. Disable all tools.

    **Answer: B.** Match boundaries to stakes: a reversible internal draft you'll review needs only light review, not heavy gating. Heavy gates (A) waste attention. A zero cap (C) or disabling tools (D) blocks a harmless task.
  </AccordionItem>

  <AccordionItem title="Q11 · An agent emailed new external contacts despite a brief saying 'never contact anyone new without checking'. What is the correct fix? (Select one)">
    A. Reword the instruction more firmly.
    B. Move the boundary into the system: enforce an approval gate on sends to new contacts so the agent cannot send without authorisation.
    C. Use a bigger model that follows instructions better.
    D. Remove the definition of done.

    **Answer: B.** The failure is a respected boundary on an irreversible action; the fix is to enforce it as a system-level approval gate. Firmer wording (A) is still just a respected boundary. A bigger model (C) can still misread. Removing the definition of done (D) is unrelated and harmful.
  </AccordionItem>

  <AccordionItem title="Q12 · A manager wants to bound the worst case of an overnight agent run over an unbounded queue. Which TWO enforced limits fit best? (Select two)">
    A. A time or step limit sized to the realistic run.
    B. A scope cap on the number of items processed.
    C. A note in the brief asking it to be efficient.
    D. A friendlier tone setting.
    E. Removing all checkpoints so it finishes faster.

    **Answer: A and B.** An enforced time/step limit (A) and a scope cap on items (B) make the worst case bounded and known for an unattended overnight run. A polite note (C) is a respected boundary that won't stop a runaway. Tone (D) is irrelevant, and removing checkpoints (E) increases risk.
  </AccordionItem>
</Accordions>

## Key takeaways

- Start every boundary decision from **reversibility**: what can this agent do that I can't take back?
- The core distinction is **respected** (instruction in the brief) versus **enforced** (system makes crossing impossible); irreversible or sensitive actions require **enforcement**.
- Put an **approval gate before** any irreversible action, surfacing enough context to decide — never after, and never on every step.
- Set **enforced spend and time/step limits** so a looping or runaway agent has a bounded, known worst case.
- The strongest **data boundary** is architectural: an agent cannot leak what it was never granted reach to (links to D3).
- **Match the boundary to the stakes**: light review for reversible low-stakes work, enforced gates and caps for irreversible or sensitive work.
- A clearer instruction never fixes a control problem — use **GATE** to turn irreversible/sensitive actions from requests into controls.
