# Human approval

_Level 03 · Workflows · sourced_

Pausing for a person to approve or correct.

## Conceptual architecture: Approval applies to a particular action.

The exact payload, destination, version, and expiry are part of the decision.

- **Proposed action:** The model drafts a concrete request
- **Validate proposal:** Schema, allowed scope, current state
- **Human review:** Show payload and consequences
- **Revise or reject:** Changed proposal needs fresh review
- **Record the outcome:** Receipt + idempotency key
- **Execute once:** Recheck approval and current state

Connections:
- Proposed action → proposal → Validate proposal
- Validate proposal → valid request → Human review
- Human review → approved → Execute once
- Execute once → receipt → Record the outcome
- Human review → changes / refusal → Revise or reject
- Revise or reject → revised proposal → Proposed action

A human clicking approve is one control, not a substitute for validation. Approval should become invalid when the reviewed action changes; execution must also handle retries and stale state.
- **Show:** What will happen, where, to whom, and whether it can be undone.
- **Bind:** Hash or version the exact proposal and scope; set an appropriate expiry.
- **Enforce:** Check again at execution, and reconcile ambiguous outcomes before retrying.

## Try this in a recipe
- [Approve the exact change before it happens](/gradient_ascent/recipes/assistant-team.md): Draft a calendar change, bind review to the exact proposal, and detect stale or repeated approvals.

## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a proposed action into a human decision and then inspect exactly what that decision permits. Compare accepting, editing, and rejecting the proposal.

**Assumptions:** Approval must identify what the person reviewed. Silence and a previous approval do not automatically cover a changed action.

**Design choices:** Put review at a meaningful commitment boundary. Low-risk routine actions can be preauthorized within a defined scope; more consequential changes may need a fresh decision.

**Request:** Let me review the weekly report and recipients before sending.

**Starting evidence:** Draft v1: Atlas delayed; Cedar unknown. Recipients: project leads. Neither has been approved.

**Action and control:** Present the exact draft, evidence, and recipient list; bind approval to that version and audience.

**Stage records (authored, not executed):**

### Review request · v1

Draft: Atlas delayed; Cedar unknown.
Recipients: project leads.
Approval: absent.
Delivery: not attempted.

What changed: The proposed content and the proposed audience are both part of the decision.

### Decision scope

A reviewer may approve, revise, or reject this version for project leads.
Approval of a draft does not imply permission to add recipients.
No response leaves delivery pending.
This is a task policy, not a universal requirement for every draft.

What changed: The reviewer sees the commitment being authorized rather than a vague request to continue.

### Review packet

Evidence: Atlas update supports delay; Cedar has no fresh update.
Open issue: Cedar's actual current status.
Allowed draft wording: unknown.
Decision requested: may this exact report go to project leads?

What changed: The reviewer can approve an honest report with an explicit gap, or request the missing information first.

### Approval workspace

Use the controls below to approve, edit, or reject.
Version and audience form the approval scope.
Approval alone does not record delivery.
Edits return the current proposal to review.

What changed: The live control state below is authoritative for this simulation; this record explains the rule.

### Audience change · illustrated diff

Before: project leads.
After: project leads + external recipient.
Unchanged: report text.
Changed: disclosure scope.
Result: prior approval no longer applies; review disclosure suitability and seek a new decision.

What changed: A recipient-only change can be material even when the words are identical.

### Your approval policy

Specify: approver, action, content, audience, and expiry if needed.
Preauthorize: low-risk routine work within a clear scope.
Re-review: material changes outside that scope.
Record separately: decision and action outcome.

What changed: Use meaningful commitment boundaries instead of asking permission for every preparatory step.

**Sample result:** Review packet v1 is ready. No send occurred. Approval applies only to v1 for project leads.

**Change something — Add an external recipient after approval:** The changed audience invalidates the old approval. Check disclosure suitability and return to review.

**Decision:** Does old approval cover an expanded recipient list?

**Answer:** No; request renewed review and approval.

**Why:** Reject or edit the draft; any change to approved content or recipients requires renewed approval. No response means no send.

**Review criteria:** A previous-versus-current diff, evidence inspection, approve/edit/reject decisions, and a clearly simulated delivery record.

**Recovery:** When content, destination, or scope changes, determine whether the existing approval still applies. Preserve the rejection or revision request so execution does not bypass it.

**Adapt it:** Use this for purchases, messages, test plans, or configuration changes. Define who can decide, what they need to see, and which changes require renewed review.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a proposed action into a human decision and then inspect exactly what that decision permits. Compare accepting, editing, and rejecting the proposal.

**Assumptions:** Approval must identify what the person reviewed. Silence and a previous approval do not automatically cover a changed action.

**Design choices:** Put review at a meaningful commitment boundary. Low-risk routine actions can be preauthorized within a defined scope; more consequential changes may need a fresh decision.

**Request:** Prepare a grocery order but ask before purchasing.

**Starting evidence:** Basket v1: $32 from Store A. Delivery address and items shown for review.

**Action and control:** Bind approval to the exact basket, total, store, and destination; this is a simulated action only.

**Stage records (authored, not executed):**

### Input record

Basket v1: $32 from Store A. Delivery address and items shown for review.

What changed: Establish the facts supplied for this version of the task.

### Design note

Put review at a meaningful commitment boundary. Low-risk routine actions can be preauthorized within a defined scope; more consequential changes may need a fresh decision.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Bind approval to the exact basket, total, store, and destination; this is a simulated action only.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Proposal: basket v1, $32, Store A, home delivery. Waiting for a decision.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Compare the approved and executed proposal, including price and substitutions.

If the result falls short:
When content, destination, or scope changes, determine whether the existing approval still applies. Preserve the rejection or revision request so execution does not bypass it.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for purchases, messages, test plans, or configuration changes. Define who can decide, what they need to see, and which changes require renewed review.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Proposal: basket v1, $32, Store A, home delivery. Waiting for a decision.

**Change something — Store substitutes an item and raises price to $39:** Revised proposal needs review. Being under a spending cap alone does not approve a substitution.

**Decision:** Does the old approval cover an altered basket?

**Answer:** No; review the new proposal.

**Why:** Approval authorizes a particular action or bounded policy, not any convenient variation.

**Review criteria:** Compare the approved and executed proposal, including price and substitutions.

**Recovery:** When content, destination, or scope changes, determine whether the existing approval still applies. Preserve the rejection or revision request so execution does not bypass it.

**Adapt it:** Use this for purchases, messages, test plans, or configuration changes. Define who can decide, what they need to see, and which changes require renewed review.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a proposed action into a human decision and then inspect exactly what that decision permits. Compare accepting, editing, and rejecting the proposal.

**Assumptions:** Approval must identify what the person reviewed. Silence and a previous approval do not automatically cover a changed action.

**Design choices:** Put review at a meaningful commitment boundary. Low-risk routine actions can be preauthorized within a defined scope; more consequential changes may need a fresh decision.

**Request:** Prepare a calibration configuration for review before it can be applied.

**Starting evidence:** Proposed set point 2.0 V; approved range 0–2.5 V; target channel A. No hardware access in this example.

**Action and control:** Review exact value, unit, target, and scope before authorizing an action; simulation does not energize equipment.

**Stage records (authored, not executed):**

### Input record

Proposed set point 2.0 V; approved range 0–2.5 V; target channel A. No hardware access in this example.

What changed: Establish the facts supplied for this version of the task.

### Design note

Put review at a meaningful commitment boundary. Low-risk routine actions can be preauthorized within a defined scope; more consequential changes may need a fresh decision.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Review exact value, unit, target, and scope before authorizing an action; simulation does not energize equipment.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Proposal: 2.0 V on channel A. Application remains blocked pending explicit approval.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Check value, unit, channel, limits, and approval identity. No simulated result proves physical safety.

If the result falls short:
When content, destination, or scope changes, determine whether the existing approval still applies. Preserve the rejection or revision request so execution does not bypass it.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for purchases, messages, test plans, or configuration changes. Define who can decide, what they need to see, and which changes require renewed review.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Proposal: 2.0 V on channel A. Application remains blocked pending explicit approval.

**Change something — Change the target to channel B after approval:** Approval for A does not transfer to B. Return the proposal to review and recheck B's limits.

**Decision:** Does approving a voltage authorize it on every channel?

**Answer:** No; target and limits belong to the approval.

**Why:** Numerical validity and authorization are separate; real hardware also needs independent protective controls.

**Review criteria:** Check value, unit, channel, limits, and approval identity. No simulated result proves physical safety.

**Recovery:** When content, destination, or scope changes, determine whether the existing approval still applies. Preserve the rejection or revision request so execution does not bypass it.

**Adapt it:** Use this for purchases, messages, test plans, or configuration changes. Define who can decide, what they need to see, and which changes require renewed review.

Human approval pauses a run before something costly, irreversible, or too uncertain to ship, and
hands that decision to a person. Anthropic frames the pause as a checkpoint inside an agent's own
loop ("Agents can then pause for human feedback at checkpoints or when encountering
blockers"[1]), but the version on this page stays at level 3: *your code* decides when
to pause, against a fixed rule. The model is never asked whether a person should look; the line
to level 4 is exactly that, a design where the model can call for review itself.

The rule that decides *when* to pause does the real work: confidence (the draft has nothing to
point to) or cost (it names a price, or an action with a consequence if it is wrong). Get the
threshold wrong either way and the gate fails at its job: too loose waves through what most
needed a look; too strict makes approving a reflex. The two kinds of mistake rarely cost the
same, and getting that threshold right on purpose is the subject of Build it, below.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 3 · Human approval. The same steps are described in the sections below._

## Practical guidance

Put the approval gate right before the step that is expensive or hard to undo, not in front of
every step. Power Automate sells the feature directly, under the heading "Streamlined approval
processes": "Create, manage, and share approval processes across your organization"[2],
pausing a flow until a person approves or rejects and letting only an approval continue it. Jules
puts two gates around the part that is genuinely hard to undo instead of one in front of
everything: you confirm a plan first, "That looks good. Continue!", and once the work is done,
"Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on
GitHub"[3].

An approval screen is only worth having if it shows enough to actually judge, not just a yes or
no button. It needs to show the thing being approved in full, the draft text or the change
itself, and what it was based on, its sources or citations, so you can check a specific claim
against a specific source rather than approving on how confident it sounds. A screen that only
asks whether to approve, with nothing to check it against, is asking you to rubber-stamp, not
review. Where missing a bad case costs far more than a false alarm, set the gate to catch more
than strictly necessary on purpose: reviewing a few extra items that turn out fine is the price
of not missing the one that does not.

Watch for the point an approval turns into a reflex. Cline's own description of the alternative
is a single setting: "Approve every step, or flip auto-approve for autopilot"[4], and
which of those you actually want is worth deciding on purpose rather than by habit. Time yourself
once: how long do you actually spend reading before you click approve, and is that long enough to
have caught a real mistake? If the honest answer is no, move the gate to the one step that truly
matters and read that one closely, or stop approving and turn on whatever the tool calls
automatic mode. A gate nobody is really reading is not a control. It is just a delay.

## Implementation details

The example checks a drafted answer against two fixed rules: no citation at all (`low_confidence`)
or a dollar figure in the text (`high_cost`), and if either trips, `run` returns a `PendingReview`
instead of a final `Answer`. `PendingReview` is a plain dataclass: the question, the full draft
text, its citations, and the reason. That is deliberately everything a reviewer needs to judge
the answer on its own merits, not just a bare yes/no: a reviewer shown only "approve this
answer?" with no sources to check it against is being asked to rubber-stamp, not review.

`resume` is a second, separate function. It takes the checkpoint, a person's decision
(`approve`, `edit` or `reject`), and, for an edit, their corrected text, and produces the final
answer. Nothing about the pause or the resume is a model decision: `_needs_review` is a plain
function of the draft's text and citations, and `resume` just branches on a string a person
supplied. A real system would serialize `PendingReview` the same way, hand it to a queue or a
ticket, and call `resume` whenever the decision comes back: hours or days later, in a different
process entirely, with nothing about the code above needing to change.

The two reasons `_needs_review` checks do not have to weigh equally. When missing a bad case
costs far more than a false alarm, bias the rule on purpose: treat anything not confidently safe
as needing a look, with the default for doubt the dangerous category, not the common one. You
will review more than you strictly need to; that is the price of the asymmetry, and the number to
watch afterward is how many of the dangerous cases in a labeled set still reached a person
unpaused. It should be none.

`examples/human_in_the_loop/run.py` (lines 26-101)

```python
LEVEL = 3
RETRIEVE_K = 4
COST_PATTERN = re.compile(r"\$\d")
DRAFT_SYSTEM = (
    "You answer questions about Halvorsen appliances using only the numbered sources below. If "
    "the sources do not answer the question, say so plainly instead of guessing. End your answer "
    "with a line starting 'Sources:' listing the citations, like 'dw300-manual#3', you used."
)

Reason = Literal["low_confidence", "high_cost"]
Decision = Literal["approve", "edit", "reject"]

@dataclass(frozen=True)
class PendingReview:
    """A paused run: everything a reviewer needs to see, and everything `resume` needs to finish
    once they decide. Every field is a plain value — this is exactly what a real system would
    persist between the pause and whenever a person actually gets to it."""

    question: str
    draft_text: str
    citations: list[str]
    reason: Reason

def _needs_review(draft_text: str, citations: list[str]) -> Reason | None:
    if not citations:
        return "low_confidence"
    if COST_PATTERN.search(draft_text):
        return "high_cost"
    return None

def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer | PendingReview:
    del embedder  # retrieval here is keyword search, not a vector index
    sections: dict[str, Section] = load_sections(corpus_dir)
    sources = [s for s, score in bm25_search(sections, question, k=RETRIEVE_K) if score > 0]
    tracer.record(kind="code", decided_by="code", title="Retrieve sources", detail=", ".join(s.cite for s in sources) or "none")

    blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
    completion = model.complete(
        [Message(role="system", content=DRAFT_SYSTEM), Message(role="user", content=f"Sources:\n\n{blocks}\n\nQuestion: {question}")],
        max_tokens=400,
    )
    tracer.record(
        kind="model", decided_by="code", title="Draft an answer", detail=completion.text[:200],
        tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )

    citations = sorted(set(CITE_RE.findall(completion.text.lower())))
    reason = _needs_review(completion.text, citations)
    tracer.record(kind="code", decided_by="code", title="Check confidence and cost thresholds", detail=f"reason={reason or 'none'}")
    if reason is None:
        return Answer(text=completion.text, citations=citations)

    tracer.record(kind="code", decided_by="code", title="Pause for human approval", detail=reason)
    return PendingReview(question=question, draft_text=completion.text, citations=citations, reason=reason)

def resume(pending: PendingReview, decision: Decision, tracer: Tracer, *, note: str = "") -> Answer:
    tracer.record(
        kind="code", decided_by="code", title="Resume from checkpoint with the reviewer's decision",
        detail=f"decision={decision}" + (f" note={note!r}" if note else ""),
    )
    if decision == "approve":
        return Answer(text=pending.draft_text, citations=pending.citations)
    if decision == "edit":
        return Answer(text=note, citations=pending.citations)
    return Answer(text="The reviewer rejected this answer; no answer is given.", citations=[])
```

Run it yourself:

`examples/human_in_the_loop/README.md` (lines 17-17)

```text
python -m examples.human_in_the_loop --model stub:scripted
```

The example's own tests stand in for the person: a small scripted "reviewer" function takes a
`PendingReview` and returns a decision, the same way a real reviewer's click would, so the pause
and the resume can both be exercised on `StubModel` with no actual person or live model involved.

If you would rather not write the checkpoint yourself, LangGraph has this built in. Its
documentation describes middleware that pauses when a model proposes an action that might need
review, waits for a decision, and saves the graph's state so the run can resume
later[5]. These are the same two halves as `run` and `resume` above, with the persistence
supplied. The decisions it names are the three this example takes plus one more: approve, edit,
reject, and answer the model directly.

`docs/THE-BENCH.md` draws this same line through an instrument's command set, written once in
`examples/common/bench.py` as `READ_ONLY_HEADERS` and `is_read_only()`: a query runs unattended
in production test, in engineering test, and in a precise measurement session alike, while
anything that sets a voltage, a current limit or an output needs a person's approval before code
will act on it, the same two-gate shape `run` and `resume` use here. The engineering recipes
[test failure triage](/gradient_ascent/recipes/test-failure-triage/),
[requirements to a test plan](/gradient_ascent/recipes/requirements-to-test-plan/),
[accuracy specs from the manual](/gradient_ascent/recipes/accuracy-specs-from-the-manual/) and
[a bring-up debug assistant](/gradient_ascent/recipes/bring-up-debug-assistant/) all turn on it:
the approval gate is code's, never the model's, and a model never produces the reported
measurement, the uncertainty, the margin or the verdict.

## When you do not need this

Try shipping without a gate first if a wrong answer costs little and is easy to notice and fix
after the fact: a gate adds latency and a person's attention, and both are wasted on an answer
nobody needed to check.

Move up to a real approval gate once being wrong is expensive, hard to undo, or the kind of
mistake nobody would notice until it was too late to matter, and pick the threshold from what
actually made past answers wrong, not a guess. This is the site's
[plan and decompose](/gradient_ascent/shapes/#plan-and-decompose) shape wherever the thing being
approved is a plan rather than an answer: nothing happens until a person signs off.

## Failure modes

### Approval fatigue

- **How to notice it:** Reviewers start approving without reading, because too many of the things they are asked to check turn out to be fine, and the gate becomes a formality rather than a control.
- **How to test for it:** Track the time between a review being shown and a decision being made. A gap that stays suspiciously short and constant, regardless of how long the draft is, is a sign the reviewer stopped actually reading.

### The threshold is tuned wrong

- **How to notice it:** Either almost everything pauses (a threshold too sensitive, breeding fatigue) or almost nothing does (a threshold too loose, so the cases that most needed a second look slip through with everything else).
- **How to test for it:** Track what share of real traffic pauses over time, and separately, sample the answers that did NOT pause and check by hand whether any of them should have.

### The checkpoint does not show enough to judge

- **How to notice it:** A reviewer is shown the draft but not what it was grounded in, so a citation that looks plausible cannot actually be checked against the source it claims to come from.
- **How to test for it:** Show a reviewer only the draft text, without the sources, and a version with the sources attached, and compare how often each version gets approved. A gap between the two says the bare draft was not enough to judge on.

### The decision never reaches resume

- **How to notice it:** A paused run sits in a queue nobody is watching, or the decision is recorded somewhere resume never reads it from, so a question a person genuinely answered never actually produces a final answer.
- **How to test for it:** Time how long a paused checkpoint sits before resume is called on it, end to end, not just how long it takes a person to click a button once they see it.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, no pause needed:** 1
- **Model calls, paused:** 1
- **Added latency when paused:** minutes to days
- **Wall time, no pause needed:** ~0.6s

**Compared with RAG (level 2), no gate.** The model cost is identical to a plain RAG call when nothing trips the threshold. The real cost of a pause is not tokens; it is the wall-clock time until a person actually looks, which can be orders of magnitude longer than the model call it is checking.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

The site's shared 60-question set grades answers, and this technique's whole point is that some
questions end without one: a run that trips `low_confidence` or `high_cost` hands back a
checkpoint for a person instead. `scripts/eval_run.py` will not score it for that reason, and
says so rather than scoring the questions that happen not to pause and calling that a number for
this technique (see `docs/EVALS.md`).

What would be measured here is the gate, in three parts. The pause rate: what share of questions
trip each threshold. Whether the right ones pause: `unanswerable` questions should pause at a
high rate, having no citation to point to, and a low pause rate on that kind means the threshold
is missing exactly the case it exists to catch. And accuracy after `resume`, scored separately
for approved and for edited answers, which is the only measurement that says whether the person
in the loop is adding anything. None of these is a token cost, and none of them can tell you how
long a real queue of paused checkpoints takes a person to clear.

## Run it

**What to monitor.** The pause rate over time, and separately, the approve/edit/reject split among decisions actually made. A pause rate that drifts with no change to the threshold code is a sign the traffic mix changed, not the rule.

**Cost at volume.** Model cost tracks question count the same as a single-call technique. The real cost that grows with volume is reviewer time: a fixed pause rate against rising traffic means a rising number of checkpoints waiting on the same number of people.

**How it fails in production.** The pause rate creeps up until approving becomes reflexive, or a queue of paused checkpoints backs up faster than anyone is clearing it, and answers that were correctly flagged for review sit unresolved rather than being wrong out loud.

**What to log.** Every pause with its reason, the full checkpoint shown to the reviewer, the decision made, who made it, and the time between the pause and the resume: an approval with no record of what was actually shown is not auditable after the fact.

## Try it

1. **Use it.** Find a tool you use that asks you to approve something before it acts (a form, an agent, an automation). Time how long you actually spend reading before you click approve. Is that enough time to have caught a real mistake?
2. **Build it.** Run python -m examples.human_in_the_loop --model stub:scripted from the repo root. The draft is a real cited answer and it pauses anyway, for high_cost, because it names $52.00; add --decision reject and the run ends with no answer given, or --decision approve and the same draft ships unchanged. Then run --model stub --question 'zzqqx frobnitz wibble', which matches nothing: that draft has no citation, so it pauses for low_confidence instead, and the two thresholds are visible one against the other.
3. **Either lane.** Take an approval step you own and write down the last three things it stopped. If you cannot name one, the threshold is either too loose to catch anything or too tight to be read.
4. **Build it.** Open READ_ONLY_HEADERS and is_read_only in examples/common/bench.py, then output_on in the same file. Name the two things a person has to approve before it will send OUTP ON, and what happens if only one of them is supplied.


## Sources

1. [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic, 2024-12-19 (accessed 2026-09-19)
2. [Power Automate](https://www.microsoft.com/en-us/power-platform/products/power-automate) — Microsoft (accessed 2026-09-19)
3. [Jules](https://jules.google/) — Google (accessed 2026-09-19)
4. [Cline](https://cline.bot/) — Cline (accessed 2026-09-19)
5. [Human-in-the-loop](https://docs.langchain.com/oss/python/langchain/human-in-the-loop) — LangChain (documentation) (accessed 2026-09-19)


Last reviewed 2026-09-19.
