The example checks a drafted answer against two fixed rules: no citation at all (low_confidence)
or a dollar figure in the text (high_cost), and if either trips, run returns a PendingReview
instead of a final Answer. PendingReview is a plain dataclass: the question, the full draft
text, its citations, and the reason. That is deliberately everything a reviewer needs to judge
the answer on its own merits, not just a bare yes/no: a reviewer shown only “approve this
answer?” with no sources to check it against is being asked to rubber-stamp, not review.
resume is a second, separate function. It takes the checkpoint, a person’s decision
(approve, edit or reject), and, for an edit, their corrected text, and produces the final
answer. Nothing about the pause or the resume is a model decision: _needs_review is a plain
function of the draft’s text and citations, and resume just branches on a string a person
supplied. A real system would serialize PendingReview the same way, hand it to a queue or a
ticket, and call resume whenever the decision comes back: hours or days later, in a different
process entirely, with nothing about the code above needing to change.
The two reasons _needs_review checks do not have to weigh equally. When missing a bad case
costs far more than a false alarm, bias the rule on purpose: treat anything not confidently safe
as needing a look, with the default for doubt the dangerous category, not the common one. You
will review more than you strictly need to; that is the price of the asymmetry, and the number to
watch afterward is how many of the dangerous cases in a labeled set still reached a person
unpaused. It should be none.
examples/human_in_the_loop/run.py · lines 26–101
LEVEL = 3
RETRIEVE_K = 4
COST_PATTERN = re.compile(r"\$\d")
DRAFT_SYSTEM = (
"You answer questions about Halvorsen appliances using only the numbered sources below. If "
"the sources do not answer the question, say so plainly instead of guessing. End your answer "
"with a line starting 'Sources:' listing the citations, like 'dw300-manual#3', you used."
)
Reason = Literal["low_confidence", "high_cost"]
Decision = Literal["approve", "edit", "reject"]
@dataclass(frozen=True)
class PendingReview:
"""A paused run: everything a reviewer needs to see, and everything `resume` needs to finish
once they decide. Every field is a plain value — this is exactly what a real system would
persist between the pause and whenever a person actually gets to it."""
question: str
draft_text: str
citations: list[str]
reason: Reason
def _needs_review(draft_text: str, citations: list[str]) -> Reason | None:
if not citations:
return "low_confidence"
if COST_PATTERN.search(draft_text):
return "high_cost"
return None
def run(
question: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer | PendingReview:
del embedder # retrieval here is keyword search, not a vector index
sections: dict[str, Section] = load_sections(corpus_dir)
sources = [s for s, score in bm25_search(sections, question, k=RETRIEVE_K) if score > 0]
tracer.record(kind="code", decided_by="code", title="Retrieve sources", detail=", ".join(s.cite for s in sources) or "none")
blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
completion = model.complete(
[Message(role="system", content=DRAFT_SYSTEM), Message(role="user", content=f"Sources:\n\n{blocks}\n\nQuestion: {question}")],
max_tokens=400,
)
tracer.record(
kind="model", decided_by="code", title="Draft an answer", detail=completion.text[:200],
tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
)
citations = sorted(set(CITE_RE.findall(completion.text.lower())))
reason = _needs_review(completion.text, citations)
tracer.record(kind="code", decided_by="code", title="Check confidence and cost thresholds", detail=f"reason={reason or 'none'}")
if reason is None:
return Answer(text=completion.text, citations=citations)
tracer.record(kind="code", decided_by="code", title="Pause for human approval", detail=reason)
return PendingReview(question=question, draft_text=completion.text, citations=citations, reason=reason)
def resume(pending: PendingReview, decision: Decision, tracer: Tracer, *, note: str = "") -> Answer:
tracer.record(
kind="code", decided_by="code", title="Resume from checkpoint with the reviewer's decision",
detail=f"decision={decision}" + (f" note={note!r}" if note else ""),
)
if decision == "approve":
return Answer(text=pending.draft_text, citations=pending.citations)
if decision == "edit":
return Answer(text=note, citations=pending.citations)
return Answer(text="The reviewer rejected this answer; no answer is given.", citations=[])
Run it yourself:
examples/human_in_the_loop/README.md · lines 17–17
python -m examples.human_in_the_loop --model stub:scripted
The example’s own tests stand in for the person: a small scripted “reviewer” function takes a
PendingReview and returns a decision, the same way a real reviewer’s click would, so the pause
and the resume can both be exercised on StubModel with no actual person or live model involved.
If you would rather not write the checkpoint yourself, LangGraph has this built in. Its
documentation describes middleware that pauses when a model proposes an action that might need
review, waits for a decision, and saves the graph’s state so the run can resume
later[5]. These are the same two halves as run and resume above, with the persistence
supplied. The decisions it names are the three this example takes plus one more: approve, edit,
reject, and answer the model directly.
docs/THE-BENCH.md draws this same line through an instrument’s command set, written once in
examples/common/bench.py as READ_ONLY_HEADERS and is_read_only(): a query runs unattended
in production test, in engineering test, and in a precise measurement session alike, while
anything that sets a voltage, a current limit or an output needs a person’s approval before code
will act on it, the same two-gate shape run and resume use here. The engineering recipes
test failure triage,
requirements to a test plan,
accuracy specs from the manual and
a bring-up debug assistant all turn on it:
the approval gate is code’s, never the model’s, and a model never produces the reported
measurement, the uncertainty, the margin or the verdict.