Level 03 · Workflows

Human approval

Pausing for a person to approve or correct.

Sourced

How it works · conceptual architecture

Approval applies to a particular action.

The exact payload, destination, version, and expiry are part of the decision.

Step / conditionInformation / relationshipReturn / repeatHighlighted box: model
Approval applies to a particular action.Proposed action → proposal → Validate proposal. Validate proposal → valid request → Human review. Human review → approved → Execute once. Execute once → receipt → Record the outcome. Human review → changes / refusal → Revise or reject. Revise or reject → revised proposal → Proposed action.proposalvalid requestapprovedreceiptchanges / refusalrevised proposalAProposed actionThe model drafts a concreterequestBValidate proposalSchema, allowed scope, currentstateCHuman reviewShow payload and consequencesDRevise or rejectChanged proposal needs freshreviewERecord the outcomeReceipt + idempotency keyFExecute onceRecheck approval and currentstate
A
Proposed action

The model drafts a concrete request

  • proposal B · Validate proposal
B
Validate proposal

Schema, allowed scope, current state

  • valid request C · Human review
C
Human review

Show payload and consequences

  • approved F · Execute once
  • changes / refusal D · Revise or reject
D
Revise or reject

Changed proposal needs fresh review

  • revised proposal A · Proposed action
E
Record the outcome

Receipt + idempotency key

    F
    Execute once

    Recheck approval and current state

    • receipt E · Record the outcome
    A human clicking approve is one control, not a substitute for validation. Approval should become invalid when the reviewed action changes; execution must also handle retries and stale state.
    The details that change the design

    Show

    What will happen, where, to whom, and whether it can be undone.

    Bind

    Hash or version the exact proposal and scope; set an appropriate expiry.

    Enforce

    Check again at execution, and reconcile ambiguous outcomes before retrying.

    CHOOSE YOUR PERSPECTIVE

    Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

    GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

    Human approval: see it in practice.

    Pausing a process for a person's scoped approval, correction, or decision.

    What you’ll walk through

    Follow a proposed action into a human decision and then inspect exactly what that decision permits. Compare accepting, editing, and rejecting the proposal.

    The task in this version

    Let me review the weekly report and recipients before sending.

    What you’ll learn to check

    A previous-versus-current diff, evidence inspection, approve/edit/reject decisions, and a clearly simulated delivery record.

    The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

    Business & team operationsAn authored case with its own evidence, changed condition, and decision.
    The task in this example

    Let me review the weekly report and recipients before sending.

    Authored case. Select any record below; nothing is sent to a model.
    FOLLOW THE EXAMPLE1 / 6
    Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
    THE VISIBLE WORKStarting evidence
    Review request · v1
    AUTHORED TEACHING RECORD · NOT A LIVE RUN
    Draft: Atlas delayed; Cedar unknown. Recipients: project leads. Approval: absent. Delivery: not attempted.

    What changed: The proposed content and the proposed audience are both part of the decision.

    WHY THIS MATTERS

    What this case assumes

    Approval must identify what the person reviewed. Silence and a previous approval do not automatically cover a changed action.

    1 / 6

    Apply this to your project

    Describe your task to your own model and use Human approval as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

    Go deeper: practical guidance, failure modes, and implementation

    Human approval pauses a run before something costly, irreversible, or too uncertain to ship, and hands that decision to a person. Anthropic frames the pause as a checkpoint inside an agent’s own loop (“Agents can then pause for human feedback at checkpoints or when encountering blockers”[1]), but the version on this page stays at level 3: your code decides when to pause, against a fixed rule. The model is never asked whether a person should look; the line to level 4 is exactly that, a design where the model can call for review itself.

    The rule that decides when to pause does the real work: confidence (the draft has nothing to point to) or cost (it names a price, or an action with a consequence if it is wrong). Get the threshold wrong either way and the gate fails at its job: too loose waves through what most needed a look; too strict makes approving a reflex. The two kinds of mistake rarely cost the same, and getting that threshold right on purpose is the subject of Build it, below.

    This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

    Optional: inspect the implementation trace

    This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

    Human approval

    Pause on a fixed threshold, hand back a checkpoint, resume once a person decides.

    Level 3 · Workflows
    QuestionQuestionRetrieve sourcesRetrieve sourcesMODELDraftDraftCheck thresholdsCheck thresholdsPERSONReviewer decidesReviewer decidesResumeResumeAnswerAnswer
    0of 1 step so far chosen by the model
    your code chose this stepthe model chose this step

    The run, step by step

    This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

    STEP 01 / 06Your code chose

    The question arrives

    "What does HLV-2205 cost?"
    0 tokens · 0 ms

    Practical guidance

    Put the approval gate right before the step that is expensive or hard to undo, not in front of every step. Power Automate sells the feature directly, under the heading “Streamlined approval processes”: “Create, manage, and share approval processes across your organization”[2], pausing a flow until a person approves or rejects and letting only an approval continue it. Jules puts two gates around the part that is genuinely hard to undo instead of one in front of everything: you confirm a plan first, “That looks good. Continue!”, and once the work is done, “Jules creates a PR of the changes. Approve the PR, merge it to your branch, and publish it on GitHub”[3].

    An approval screen is only worth having if it shows enough to actually judge, not just a yes or no button. It needs to show the thing being approved in full, the draft text or the change itself, and what it was based on, its sources or citations, so you can check a specific claim against a specific source rather than approving on how confident it sounds. A screen that only asks whether to approve, with nothing to check it against, is asking you to rubber-stamp, not review. Where missing a bad case costs far more than a false alarm, set the gate to catch more than strictly necessary on purpose: reviewing a few extra items that turn out fine is the price of not missing the one that does not.

    Watch for the point an approval turns into a reflex. Cline’s own description of the alternative is a single setting: “Approve every step, or flip auto-approve for autopilot”[4], and which of those you actually want is worth deciding on purpose rather than by habit. Time yourself once: how long do you actually spend reading before you click approve, and is that long enough to have caught a real mistake? If the honest answer is no, move the gate to the one step that truly matters and read that one closely, or stop approving and turn on whatever the tool calls automatic mode. A gate nobody is really reading is not a control. It is just a delay.

    Implementation details

    The example checks a drafted answer against two fixed rules: no citation at all (low_confidence) or a dollar figure in the text (high_cost), and if either trips, run returns a PendingReview instead of a final Answer. PendingReview is a plain dataclass: the question, the full draft text, its citations, and the reason. That is deliberately everything a reviewer needs to judge the answer on its own merits, not just a bare yes/no: a reviewer shown only “approve this answer?” with no sources to check it against is being asked to rubber-stamp, not review.

    resume is a second, separate function. It takes the checkpoint, a person’s decision (approve, edit or reject), and, for an edit, their corrected text, and produces the final answer. Nothing about the pause or the resume is a model decision: _needs_review is a plain function of the draft’s text and citations, and resume just branches on a string a person supplied. A real system would serialize PendingReview the same way, hand it to a queue or a ticket, and call resume whenever the decision comes back: hours or days later, in a different process entirely, with nothing about the code above needing to change.

    The two reasons _needs_review checks do not have to weigh equally. When missing a bad case costs far more than a false alarm, bias the rule on purpose: treat anything not confidently safe as needing a look, with the default for doubt the dangerous category, not the common one. You will review more than you strictly need to; that is the price of the asymmetry, and the number to watch afterward is how many of the dangerous cases in a labeled set still reached a person unpaused. It should be none.

    examples/human_in_the_loop/run.py · lines 26–101
    LEVEL = 3
    RETRIEVE_K = 4
    COST_PATTERN = re.compile(r"\$\d")
    DRAFT_SYSTEM = (
        "You answer questions about Halvorsen appliances using only the numbered sources below. If "
        "the sources do not answer the question, say so plainly instead of guessing. End your answer "
        "with a line starting 'Sources:' listing the citations, like 'dw300-manual#3', you used."
    )
    
    Reason = Literal["low_confidence", "high_cost"]
    Decision = Literal["approve", "edit", "reject"]
    
    
    @dataclass(frozen=True)
    class PendingReview:
        """A paused run: everything a reviewer needs to see, and everything `resume` needs to finish
        once they decide. Every field is a plain value — this is exactly what a real system would
        persist between the pause and whenever a person actually gets to it."""
    
        question: str
        draft_text: str
        citations: list[str]
        reason: Reason
    
    
    def _needs_review(draft_text: str, citations: list[str]) -> Reason | None:
        if not citations:
            return "low_confidence"
        if COST_PATTERN.search(draft_text):
            return "high_cost"
        return None
    
    
    def run(
        question: str,
        model: Model,
        embedder: Embedder | None,
        tracer: Tracer,
        *,
        corpus_dir: Path = DEFAULT_CORPUS_DIR,
    ) -> Answer | PendingReview:
        del embedder  # retrieval here is keyword search, not a vector index
        sections: dict[str, Section] = load_sections(corpus_dir)
        sources = [s for s, score in bm25_search(sections, question, k=RETRIEVE_K) if score > 0]
        tracer.record(kind="code", decided_by="code", title="Retrieve sources", detail=", ".join(s.cite for s in sources) or "none")
    
        blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
        completion = model.complete(
            [Message(role="system", content=DRAFT_SYSTEM), Message(role="user", content=f"Sources:\n\n{blocks}\n\nQuestion: {question}")],
            max_tokens=400,
        )
        tracer.record(
            kind="model", decided_by="code", title="Draft an answer", detail=completion.text[:200],
            tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
        )
    
        citations = sorted(set(CITE_RE.findall(completion.text.lower())))
        reason = _needs_review(completion.text, citations)
        tracer.record(kind="code", decided_by="code", title="Check confidence and cost thresholds", detail=f"reason={reason or 'none'}")
        if reason is None:
            return Answer(text=completion.text, citations=citations)
    
        tracer.record(kind="code", decided_by="code", title="Pause for human approval", detail=reason)
        return PendingReview(question=question, draft_text=completion.text, citations=citations, reason=reason)
    
    
    def resume(pending: PendingReview, decision: Decision, tracer: Tracer, *, note: str = "") -> Answer:
        tracer.record(
            kind="code", decided_by="code", title="Resume from checkpoint with the reviewer's decision",
            detail=f"decision={decision}" + (f" note={note!r}" if note else ""),
        )
        if decision == "approve":
            return Answer(text=pending.draft_text, citations=pending.citations)
        if decision == "edit":
            return Answer(text=note, citations=pending.citations)
        return Answer(text="The reviewer rejected this answer; no answer is given.", citations=[])

    Run it yourself:

    examples/human_in_the_loop/README.md · lines 17–17
    python -m examples.human_in_the_loop --model stub:scripted

    The example’s own tests stand in for the person: a small scripted “reviewer” function takes a PendingReview and returns a decision, the same way a real reviewer’s click would, so the pause and the resume can both be exercised on StubModel with no actual person or live model involved.

    If you would rather not write the checkpoint yourself, LangGraph has this built in. Its documentation describes middleware that pauses when a model proposes an action that might need review, waits for a decision, and saves the graph’s state so the run can resume later[5]. These are the same two halves as run and resume above, with the persistence supplied. The decisions it names are the three this example takes plus one more: approve, edit, reject, and answer the model directly.

    docs/THE-BENCH.md draws this same line through an instrument’s command set, written once in examples/common/bench.py as READ_ONLY_HEADERS and is_read_only(): a query runs unattended in production test, in engineering test, and in a precise measurement session alike, while anything that sets a voltage, a current limit or an output needs a person’s approval before code will act on it, the same two-gate shape run and resume use here. The engineering recipes test failure triage, requirements to a test plan, accuracy specs from the manual and a bring-up debug assistant all turn on it: the approval gate is code’s, never the model’s, and a model never produces the reported measurement, the uncertainty, the margin or the verdict.

    When you do not need this

    Try shipping without a gate first if a wrong answer costs little and is easy to notice and fix after the fact: a gate adds latency and a person’s attention, and both are wasted on an answer nobody needed to check.

    Move up to a real approval gate once being wrong is expensive, hard to undo, or the kind of mistake nobody would notice until it was too late to matter, and pick the threshold from what actually made past answers wrong, not a guess. This is the site’s plan and decompose shape wherever the thing being approved is a plan rather than an answer: nothing happens until a person signs off.

    Failure modes

    Approval fatigue

    How to notice it
    Reviewers start approving without reading, because too many of the things they are asked to check turn out to be fine, and the gate becomes a formality rather than a control.
    How to test for it
    Track the time between a review being shown and a decision being made. A gap that stays suspiciously short and constant, regardless of how long the draft is, is a sign the reviewer stopped actually reading.

    The threshold is tuned wrong

    How to notice it
    Either almost everything pauses (a threshold too sensitive, breeding fatigue) or almost nothing does (a threshold too loose, so the cases that most needed a second look slip through with everything else).
    How to test for it
    Track what share of real traffic pauses over time, and separately, sample the answers that did NOT pause and check by hand whether any of them should have.

    The checkpoint does not show enough to judge

    How to notice it
    A reviewer is shown the draft but not what it was grounded in, so a citation that looks plausible cannot actually be checked against the source it claims to come from.
    How to test for it
    Show a reviewer only the draft text, without the sources, and a version with the sources attached, and compare how often each version gets approved. A gap between the two says the bare draft was not enough to judge on.

    The decision never reaches resume

    How to notice it
    A paused run sits in a queue nobody is watching, or the decision is recorded somewhere resume never reads it from, so a question a person genuinely answered never actually produces a final answer.
    How to test for it
    Time how long a paused checkpoint sits before resume is called on it, end to end, not just how long it takes a person to click a button once they see it.

    Cost and latency

    Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

    1Model calls, no pause needed
    1Model calls, paused
    minutes to daysAdded latency when paused
    ~0.6sWall time, no pause needed
    Compared with RAG (level 2), no gateThe model cost is identical to a plain RAG call when nothing trips the threshold. The real cost of a pause is not tokens; it is the wall-clock time until a person actually looks, which can be orders of magnitude longer than the model call it is checking.

    How to Evaluate It

    60 questionslookupmulti-hopnumericunanswerableconflicting sources

    The site’s shared 60-question set grades answers, and this technique’s whole point is that some questions end without one: a run that trips low_confidence or high_cost hands back a checkpoint for a person instead. scripts/eval_run.py will not score it for that reason, and says so rather than scoring the questions that happen not to pause and calling that a number for this technique (see docs/EVALS.md).

    What would be measured here is the gate, in three parts. The pause rate: what share of questions trip each threshold. Whether the right ones pause: unanswerable questions should pause at a high rate, having no citation to point to, and a low pause rate on that kind means the threshold is missing exactly the case it exists to catch. And accuracy after resume, scored separately for approved and for edited answers, which is the only measurement that says whether the person in the loop is adding anything. None of these is a token cost, and none of them can tell you how long a real queue of paused checkpoints takes a person to clear.

    Run it

    What to monitor

    The pause rate over time, and separately, the approve/edit/reject split among decisions actually made. A pause rate that drifts with no change to the threshold code is a sign the traffic mix changed, not the rule.

    Cost at volume

    Model cost tracks question count the same as a single-call technique. The real cost that grows with volume is reviewer time: a fixed pause rate against rising traffic means a rising number of checkpoints waiting on the same number of people.

    How it fails in production

    The pause rate creeps up until approving becomes reflexive, or a queue of paused checkpoints backs up faster than anyone is clearing it, and answers that were correctly flagged for review sit unresolved rather than being wrong out loud.

    What to log

    Every pause with its reason, the full checkpoint shown to the reviewer, the decision made, who made it, and the time between the pause and the resume: an approval with no record of what was actually shown is not auditable after the fact.

    Try it

    1. Use it

      Find a tool you use that asks you to approve something before it acts (a form, an agent, an automation). Time how long you actually spend reading before you click approve. Is that enough time to have caught a real mistake?

    2. Build it

      Run python -m examples.human_in_the_loop --model stub:scripted from the repo root. The draft is a real cited answer and it pauses anyway, for high_cost, because it names $52.00; add --decision reject and the run ends with no answer given, or --decision approve and the same draft ships unchanged. Then run --model stub --question 'zzqqx frobnitz wibble', which matches nothing: that draft has no citation, so it pauses for low_confidence instead, and the two thresholds are visible one against the other.

    3. Either lane

      Take an approval step you own and write down the last three things it stopped. If you cannot name one, the threshold is either too loose to catch anything or too tight to be read.

    4. Build it

      Open READ_ONLY_HEADERS and is_read_only in examples/common/bench.py, then output_on in the same file. Name the two things a person has to approve before it will send OUTP ON, and what happens if only one of them is supplied.

    How it connects

    Before, after and instead of this

    Decoded in

    Optional: products, tools, and models

    5 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

    In practice

    Approve a proposed refund

    The system prepares the refund, pauses for a person to review it, and proceeds only after approval.

    Out there

    Named products, tools and models

    Products3
    • ClineCline · open-source coding agent
    • JulesGoogle · coding agent
    • Power AutomateMicrosoft · automation service
    Tools2
    • Deep AgentsLangChain · agent harness
    • LangGraphLangChain · graph framework

    Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

    Where this comes from

    Primary sources

    1. Building effective agents · Anthropic, 12/19/2024 (accessed 09/19/2026)
    2. Power Automate · Microsoft (accessed 09/19/2026)
    3. Jules · Google (accessed 09/19/2026)
    4. Cline · Cline (accessed 09/19/2026)
    5. Human-in-the-loop · LangChain (documentation) (accessed 09/19/2026)

    Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page