Everything around the model in an agent: the loop, tools, context handling, permissions, caps and sandbox.
Sourced
How it works · conceptual architecture
The harness makes the loop executable.
Every model request passes through application controls before it affects the world.
Step / conditionInformation / relationshipReturn / repeatHighlighted box: model
A
Goal + context
Task, instructions, selected history
context →B · Model decision
B
Model decision
Request a tool or return an answer
tool request →C · Execution gate
final answer →F · Finish or hand back
C
Execution gate
Arguments, permissions, budgets
allowed →D · Run allowed tool
cannot proceed →G · Pause or refuse
D
Run allowed tool
Bounded operation in the environment
observation →E · Observe the result
E
Observe the result
Return output or a useful error
next decision →B · Model decision
F
Finish or hand back
Return work, evidence, and gaps
G
Pause or refuse
Approval needed, denied, or capped
Reasoning helps the model choose useful actions. The loop supplies feedback; the harness supplies execution, state, and enforced limits. None of those makes the answer automatically correct.The details that change the design
Control
Tool output is evidence, not permission to take another action.
Stopping
Finish, ask for help, or stop at a step, time, or cost limit.
Verification
Inspect the environment and the final artifact, not just the model’s account of its work.
Apply this to your project
Describe your task to your own model and use The agent harness as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
Go deeper: practical guidance, failure modes, and implementation
An agent harness is everything around the model in an agent: the loop that calls it, the tool
definitions it is shown and the code that runs them, what goes into its next request, whether an
action needs approval, the caps on steps and tokens, the sandbox, and what gets logged. None of that
is the model. Anthropic defines an agent in one sentence: systems “where LLMs dynamically direct
their own processes and tool usage, maintaining control over how they accomplish tasks”[1],
and almost everything a builder builds sits outside it.
That is why the same model behaves very differently in a different harness: change the step cap, the
allowlist or the context policy and a run finishes, fails, or runs up a bill doing neither, with
nothing about the model different. Anthropic’s Claude Code team calls this loop engineering and
defines a loop as “agents repeating cycles of work until a stop condition is met”[4].
Level 5 is where the harness first has real decisions to bound: the model decides both the action
and when to stop: see single agent for the loop
itself.
The harness is also where several supporting topics meet:
guardrails check inputs, outputs, and proposed actions;
human approval handles actions that need a person’s decision;
context engineering shapes the next request;
and observability records what happened.
Evaluation checks the resulting system, while
cost controls bound its work.
These are design choices around the loop, not capabilities guaranteed by the word “harness.”
Guardrail checks complement permission boundaries and sandboxing; they do not replace them.
This page is sourced, not measured: what the harnesses below do comes from their makers’ own
documentation, and no run under one has been recorded and scored here. It is illustrated.
Worked example: a test automation framework
A team uses a shared Python test automation framework. Each project represents a device under
test (DUT), follows the same project-file conventions, and uses the same or similar instruments.
The framework provides measurement methods, unit conversions, CSV export, and instrument drivers.
The drivers use SCPI, but project code treats instruments as black boxes through the framework’s
interfaces. Python files contain project code; YAML/JSON files hold configuration such as instrument
settings, test parameters, limits, and sequences, according to the framework’s conventions.
Claude Code supplies the agent harness; the shared framework supplies the domain interfaces
and conventions. Claude Code provides the model’s read, edit, and command-execution loop.[8]
The agent uses that loop to create and refine a DUT project. The user reviews the files, performs
hardware testing, and returns logs and observations. The agent does not operate the instruments.
This is a specified workflow, not a recorded implementation or hardware validation result.
Simulation is not an established capability of this framework. Mocked function outputs could be
considered later, but are not assumed here.
The walkthrough below illustrates the agent workflow with scripted responses and sample artifacts.
It does not simulate instrument physics, execute project code, or call a model. Use Watch it
to follow the task, Change something to explore a missing requirement or new helper, and
Try a decision to check your understanding. The detailed reference follows the walkthrough.
CHOOSE YOUR PERSPECTIVE
Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.
GUIDED WORKED EXAMPLE Scripted simulation · no model calls or hardware access
From DUT brief to reviewable project.
Follow one task. See what the agent proposes, what the harness controls, and where you decide.
What you’ll walk through
Imagine your team already has a Python test framework and several past device projects. You need a project for a new device under test (DUT), with different requirements but familiar instruments and conventions. This walkthrough follows a coding agent from reading that context to handing over generated files for a person to review and test.
The task in this version
Use reusable instructions in CLAUDE.md, the new DUT brief, framework documentation, and a suitable reference project to prepare Python tests, configuration, and Markdown documentation.
What you’ll learn to check
Watch the plan become a scoped implementation, see a missing requirement or proposed helper trigger a decision, and distinguish generated files from evidence that the real measurements work.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
How this applies beyond this test framework
The reusable pattern is context → proposed work → execution within authority → evidence → handoff. In this team, shared-framework edits, new helpers, and instrument access require separate authorization. Another project can preauthorize routine edits or safe checks. Choose boundaries around ownership, reversibility, and consequences rather than copying every restriction.
Engineering & technical workCreate a DUT project within the shared framework’s conventions and approval boundaries.
Prefilled English request. You can edit it; this walkthrough always demonstrates the same scripted workflow.
Always in effectProject-only scopeFramework changes need authorizationNo live instrument accessIntended controls, illustrated here.
THE VISIBLE WORKStart with a task
01 → 06
A project takes shape, one decision at a time.
Send the request above. Then follow the agent’s work and make the approval decision yourself.
Brief → Plan → Approval → Files → Evidence
Go deeper: instructions, approval boundaries, code, and evaluation
Start with an ordinary request
The user supplies two starting documents. CLAUDE.md holds reusable framework instructions;
DUT_BRIEF.md describes this particular DUT. Establish the framework instructions once and
maintain them as conventions evolve. Write a new brief for each DUT.
Document
Who provides it
What it contains
CLAUDE.md
User or framework maintainer; reused across DUT projects
The agent also needs access to framework documentation, framework source, previous projects,
and the standard template. These remain the reference material; the two starting files do not
replace them. Point to their actual locations in CLAUDE.md, and ask the agent to read
DUT_BRIEF.md when starting the task. The brief, plan, and status filenames are conventions for
this example, not special files automatically understood by every agent.
Configure permissions separately. Instructions in Markdown describe the boundaries; they do
not themselves make the framework read-only or prevent instrument access.
With those files in place, the user can give this request:
“Read CLAUDE.md and DUT_BRIEF.md, then create a project for this new DUT using our shared Python test framework. Use the closest past
project where one is suitable; otherwise use the standard template. Here is my description of how
this DUT differs. Read the framework documentation and reuse its existing tools and instrument
interfaces. Ask me about missing requirements before proceeding. Show me your proposed files and
changes in PROJECT_PLAN.md, and wait for my approval before generating the project. Produce Python code, YAML/JSON
configuration, and Markdown documentation, including PROJECT_STATUS.md, for my review. Obtain separate approval before creating any new project-local tool. Do not modify the shared framework or
connect to instruments.”
What belongs where
Part
Role in this example
Claude
Interprets the DUT differences, asks questions, proposes a plan, and drafts revisions.
Claude Code
Provides context management, file editing, command execution, and permission controls around the model.
Past projects and documentation
Supply the closest starting point, APIs, file conventions, and established patterns. A template is the fallback.
Shared Python framework
Provides reusable tools, measurement methods, unit conversion, CSV export, and instrument interfaces backed by SCPI drivers.
New DUT project
Contains the generated Python files, YAML/JSON configuration, and Markdown documentation.
Non-hardware checks
Check Python syntax and YAML/JSON validity without connecting to instruments.
User
Approves the plan, reviews the files, tests with real instruments, and supplies feedback.
The agent can generate code that imports existing framework functions. Each function does not
need its own model-tool definition. The project’s use of an instrument API does not authorize
the agent to execute it against connected equipment.
The feedback loop in practice
Understand: read the framework documentation, candidate reference projects, and the written
description of this DUT’s differences. Ask about missing requirements before filling them in.
Plan and wait: identify the closest suitable project or the standard template. Propose the
files, intended changes, framework tools to reuse, and non-hardware checks. Obtain the user’s
approval before generating project files and code.
Generate: create the approved Python files, YAML/JSON configuration, and Markdown documents
in the new project’s folder. Reuse the framework’s existing tools wherever applicable.
Check without hardware: run approved syntax and configuration checks that cannot connect
to instruments. Do not import or execute project setup code unless its lack of hardware access
is established. Report the checks performed and their results; do not label the project as
hardware-tested.
Hand over: the user reviews the project, runs it on real instruments, and supplies logs,
results, and observations. The documentation distinguishes generated work from verified behavior.
Refine: use that feedback to revise the project within the approved scope. Ask about newly
missing requirements, and seek approval for changes that cross the boundaries below.
The agent automates project creation and revision. Hardware testing remains a human-controlled
step. A passing syntax or configuration check does not establish that a measurement is correct.
Approval boundaries and hard controls
The shared framework must not be modified without explicit authorization. If a new framework
feature is absolutely required, the agent must explain the requirement, why existing capabilities
cannot satisfy it, and the proposed change, then wait for the user’s decision. Plan approval for a
DUT project is not blanket permission to modify the framework.
A new project-local tool also requires approval. Before creating one, explain the gap, which
framework tools were considered, and why a new tool is necessary. Placing a helper in the project
folder does not bypass this rule. Missing requirements must be asked about first, rather than
silently guessed or left as unapproved TODOs.
These are required boundaries, but a written instruction alone is not a hard enforcement mechanism.
Read-only access to the shared framework is a proposed control whose availability still needs to
be confirmed. The intended setup gives the agent write access only to the approved DUT project,
withholds live instrument access, and keeps permission controls outside files it can rewrite.
Any authorized framework change would need a separately scoped exception. This page does not claim
those controls are already implemented.
Markdown documentation to hand over
Capabilities and created files: supported tests, what was generated, and where each part lives.
Differences from the reference: what changed for this DUT and why.
Configuration and operation: parameters, limits, expected instruments, connections, setup,
cleanup, and instructions for the user to run the project through the framework.
Requirements checklist: each requested test mapped to its code, configuration, and validation status.
Validation record: checks the agent actually ran, followed by hardware results the user supplies.
Open questions, limitations, and approvals: unresolved issues and any requested tool or framework changes.
How the concepts fit together
The harness coordinates these concepts during one task: turn a DUT brief into a reviewable
project using an existing framework. Some parts describe what the model sees, some decide
what may happen, and others establish what actually happened. The controls below describe the
intended setup; they are not a claim that custom checks or restrictions already exist.
CLAUDE.md sets reusable rules; the request and DUT_BRIEF.md define the task.
Tell the agent to reuse framework functions, ask about unknowns, and deliver reviewable files. Instructions express policy; they do not enforce permissions.
Select the relevant API documentation, closest past project, DUT differences, approved plan, and latest feedback for the next model call.
Keep current requirements and approvals available as the conversation grows. Old project values are reference material, not automatically valid limits for this DUT.
Read, propose, wait for approval, edit, check, inspect results, and revise.
The model chooses its next action within the allowed scope; the harness executes permitted actions and returns their results. Pause for missing requirements or an approval decision.
File reads, edits, and approved non-hardware checks are actions available to the coding agent. Generated Python calls the framework APIs later when the user runs it.
Distinguish an agent tool from a Python function used by the resulting project. Writing an instrument call does not grant permission to execute it.
Proposed checks inspect intended actions and generated files for disallowed paths, direct instrument access, missing required configuration, or unapproved helpers.
Reject a prohibited action or flag work for correction. Syntax and schema checks can be deterministic; judging whether an existing tool meets a requirement may still need human review. These checks must be implemented and tested.
Permissions and sandboxing
Configure project-only writes, protected framework files, and an environment without live instrument access.
Bound what executed code can actually touch, even if the model proposes otherwise. Read-only framework access remains to be confirmed; a path check alone is not a complete sandbox.
Retain file diffs, commands, check results, approval decisions, and reasons for blocked actions; summarize progress in PROJECT_STATUS.md.
Explain why the run changed a file, stopped, or failed. The status document is a readable summary, not a substitute for the underlying execution record.
Independently compare the output with the DUT requirements, framework conventions, approval record, and reported validation status.
Assess the agent’s work. User-run hardware tests assess the resulting measurement behavior; passing syntax checks establishes neither of these on its own.
Save the approved plan, project status, and user feedback; explicitly load them when resuming.
Carry decisions between sessions without assuming the model remembers them. Saved files only help when their relevant contents reach the next request.
One proposed helper, several different controls
Suppose the agent believes it needs a new unit-conversion helper for the DUT:
Context and tools: it reads the framework’s existing conversion API and relevant past code.
If those already meet the requirement, it uses them in the generated project.
Guardrail and approval: if it proposes a new helper, a configured policy check should pause
that creation until a specific approval exists. The agent explains the gap and asks the user.
Without an implemented check, this remains an instruction the agent is expected to follow.
Permissions: approval for a helper in the project does not unlock the shared framework or
instrument access. A framework change would require its own authorization and scoped access.
Execution and observation: after approval, it writes the helper within scope, runs only
approved non-hardware checks, and records the change, approval, and actual results.
Evaluation and feedback: the reviewer checks the conversion against the agreed requirement.
The user performs any required hardware validation and returns findings. The agent revises
within scope or stops and reports the next decision it needs.
The guardrail checks the proposal; the person authorizes an exception; permissions constrain
execution; logs record the outcome; evaluation judges whether the result meets the requirement.
The harness brings these together around the model’s repeated calls.
Relating this to the code and run below
The runnable demonstration below uses a document lookup task, not this DUT framework. Its
ContextPolicy corresponds to selecting the documentation and decisions the model sees;
ToolRegistry corresponds to the actions the coding agent can call; a Hook illustrates checking
a proposed action before execution; and the caps bound the loop. A veto hook alone is not a human
approval workflow: that also needs a pause, a recorded decision, and a way to resume within scope.
The trace illustrates observability. The DUT diagram describes how these responsibilities would
apply to your project workflow, without claiming the demonstration implements its controls.
You do not need every technique on the site to start this workflow. Reading repository files does
not by itself establish a RAG system, a Markdown instruction file is not automatically a packaged
skill, and using Python APIs does not require MCP. Those are separate choices if retrieval,
reusable procedures, or external tool connections become necessary.
Optional: inspect the implementation trace
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
The agent harness
The model picks each action; the harness around it decides what the model can see, what may run, and when to stop.
Level 5 · Agent loops
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step
The run, step by step
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
STEP 01 / 07Your code chose
The question arrives
"What does the DW-300's drain pump cost, and
how long is it under warranty?"
0 tokens · 0 ms
Practical guidance
If you use a chat app and will never run an agent, skip this page. A harness is the code wrapped
around the model, written by whoever built the agent product, and there is no box for you to type
in. The pages that are yours are coding agents and
always-on assistants.
If you do operate an agent product, one thing here earns your time: when an agent behaves badly,
the fix is usually a setting rather than a better prompt. Four settings, and what each one looks
like when it is the cause.
What it remembers. Anthropic’s Claude Agent SDK documents automatic compaction firing as a
long session grows: “When the context window approaches its limit, the SDK automatically compacts
the conversation: it summarizes older history to free space, keeping your most recent exchanges
and key decisions intact”[2]. Anthropic’s engineering blog describes tool result
clearing, a lighter version of the same idea, as one of the “safest lightest touch forms of
compaction”[3]. An agent that dropped your constraint two hours into a session did not
ignore it; it summarized it away. Restate the constraint in your next message rather than starting
the whole task again.
What it may run without asking. OpenAI’s Codex documentation states the split: “Sandboxing and
approvals are different controls that work together. The sandbox defines technical boundaries. The
approval policy decides when the agent must stop and ask before crossing them”[6],
enforced on macOS “using the built-in Seatbelt framework”[6]. Between asking every time
and never asking, a policy can “keep specific approval prompt categories interactive while
automatically rejecting others”[7]. Start on the setting that asks, and loosen one
category at a time once you have watched what that category actually does.
What it costs before it stops. A cap on steps or spend is a harness decision, not a model one,
and a run that hits one usually ends mid-task with no error: see
single agent.
What a session hands on. A subagent may explore at length and return “only a condensed,
distilled summary of its work”[3] to the harness that spawned it. That is why an agent’s
account of what it did can be thinner than what it did, and why
memory is a separate setting from the summary.
Read those four in your own product’s documentation before concluding a rough session was the
model’s fault.
Implementation details
The document lookup demonstration is the same loop single agent runs —
act, check a cap, repeat: rebuilt so four moving parts are arguments to run instead of fixed in
the function body: a ToolRegistry (the definitions the model is shown, and the allowed set
checked before any of them runs: the split
function calling makes for one call, made
reusable), a ContextPolicy (a function from the growing message history to whatever the next
request sends: context engineering applied
inside the loop rather than once before it), a Hook (a chance to veto a call the model already
chose, before the registry runs it: a silent version of what
human approval does out loud), and the step and
token caps.
trim_to_budget is the context policy worth reading closely. It keeps every message except tool
results, and keeps only as many of the most recent tool results as fit under a token budget,
replacing older ones with a short placeholder rather than deleting them silently: a small version
of what Anthropic calls tool result clearing[3]:
examples/agent_harness/run.py · lines 75–112
def trim_to_budget(budget_tokens: int) -> ContextPolicy:
"""The tight policy: keeps every non-tool-result message, and as many of the most recent
tool results as fit under `budget_tokens`, dropping older ones first. Real context policies
trim the same way -- see this page's Use it lane for how Anthropic describes tool result
clearing and compaction -- this one trims by a plain token count to keep the point readable
in a few lines.
A tool result is recognized by its role, `tool`, never by how its text opens: the question is
a user message, so a policy that matched on text could drop a question that happened to begin
"Result of ...", leaving the model answering something it can no longer see. A trimmed result
keeps its call id, because every tool call in the history still needs an answer."""
def policy(messages: list[Message]) -> list[Message]:
result_idx = [i for i, m in enumerate(messages) if _is_tool_result(m)]
kept: set[int] = set()
used = 0
for i in reversed(result_idx):
cost = count_tokens(content_text(messages[i].content))
if used + cost > budget_tokens:
break
used += cost
kept.add(i)
out = []
for i, m in enumerate(messages):
if i in result_idx and i not in kept:
out.append(
Message(
role="tool",
content="[earlier tool result trimmed by the context policy]",
tool_call_id=m.tool_call_id,
tool_name=m.tool_name,
)
)
else:
out.append(m)
return out
return policy
It runs fresh on every model call, not once at the start, which is what lets the two runs below
diverge partway through instead of only at the first prompt. It also leaves the opening request
alone: it recognizes a tool result by how the text opens, so without that guard a question starting
the same way would be trimmed and the model would be answering something it could no longer see.
deny_after is the hook worth reading next, three lines, and the point is that it runs after the
model has already decided:
examples/agent_harness/run.py · lines 129–142
def deny_after(allowed_calls: int) -> Hook:
"""A hook for demonstration and testing: allows the first `allowed_calls` tool calls the
model attempts, vetoes every one after. A real hook would read the call's own name and
arguments; this one only counts, to keep the point -- a hook can block an action the model
already decided to take -- in three lines."""
seen = {"n": 0}
def hook(call: ToolCall) -> tuple[bool, str]:
seen["n"] += 1
if seen["n"] > allowed_calls:
return False, f"tool budget of {allowed_calls} call(s) already spent"
return True, ""
return hook
A real hook would look at the call’s own name and arguments instead of just counting; running one
inside a sandbox that actually isolates what a tool may touch, rather than a check like this one,
is code execution‘s territory, and what gets loaded
into the model’s instructions in the first place (which skill, not just which tool) is
skills’.
run is the loop these parts plug into: ask the model through whatever the context policy
currently allows it to see, and if it calls tools, check each one against the hook and the
registry’s allowlist before running it, record what happened, and go around again until the model
stops or a cap does.
examples/agent_harness/run.py · lines 145–197
def run(
question: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
corpus_dir=DEFAULT_CORPUS_DIR,
registry: ToolRegistry = DEFAULT_REGISTRY,
context_policy: ContextPolicy = keep_everything,
hook: Hook = allow_everything,
max_steps: int = MAX_STEPS,
max_tokens: int = MAX_TOKENS,
) -> Answer:
del embedder # this harness retrieves through its tools, not a vector index
sections = load_sections(corpus_dir)
messages = [Message(role="system", content=SYSTEM), Message(role="user", content=question)]
citations: list[str] = []
tokens_used = 0
for _ in range(max_steps):
completion = model.complete(context_policy(messages), tools=registry.definitions, max_tokens=400)
tokens_used += completion.tokens_in + completion.tokens_out
if not completion.tool_calls:
record_completion(tracer, decided_by="model", title="Model stops and answers", completion=completion)
return Answer.from_text(completion.text, retrieved_sources=citations)
calls_desc = ", ".join(f"{c.name}({c.arguments})" for c in completion.tool_calls)
record_completion(tracer, decided_by="model", title="Model picks an action", completion=completion, detail=calls_desc)
turn, calls = assistant_turn(completion, len(messages))
messages.append(turn)
for call in calls:
allowed, reason = hook(call)
if not allowed:
tracer.record(kind="code", decided_by="code", title="Hook vetoes the call", detail=reason)
messages.append(tool_result(call, f"Denied: {reason}"))
continue
if call.name not in registry.allowed:
result_text, cites = toolkit.unknown_tool(call.name)
else:
result_text, cites = registry.call(call, sections)
citations.extend(cites)
tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])
messages.append(tool_result(call, result_text))
if tokens_used >= max_tokens:
reason = f"token budget reached: {tokens_used} >= {max_tokens}"
final = force_final(context_policy(messages), model, tracer, reason=reason, max_tokens=400)
return Answer.from_text(final.text, retrieved_sources=citations)
final = force_final(context_policy(messages), model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
return Answer.from_text(final.text, retrieved_sources=citations)
Every tool call, its arguments, and the decision to stop are decided_by: "model"; running a tool,
a hook’s veto, and forcing a final answer when a cap is reached are always decided_by: "code":
the same split single_agent’s example makes, with two more kinds of code-decided step than that
one has.
tests/test_example_agent_harness.py runs the same scripted model twice with only the context
policy changed, and the two runs answer differently: confidently citing the warranty term under a
generous policy, saying it could not confirm the term under a tight one. Nothing about the model’s
own logic changed between the two runs; only what the harness let it see did. Read that for what it
is: a scripted stand-in, written to answer from whatever the harness left in front of it, so what
the test proves is the mechanism, not a measurement of how much a real model’s answers move. The
size of that effect is what the eval below is for, and no run of it exists yet. The same file scripts a hook that vetoes a call the
model already committed to, and checks the veto shows up in the trace as the harness’s own
decision, and a tool the registry advertises but will not run, and checks it fails exactly the way
an unknown tool does.
Try single agent first if you have not seen the basic
loop yet: this page assumes you have, and is about what surrounds it, not the loop itself.
Move to thinking about the harness deliberately once an agent runs past a one-off demo: choosing
the caps, the allowlist, the context policy and the approval settings on purpose is what turns a
loop that happens to work into a system somebody can operate and debug: see
safety, privacy and governance for testing one before
trusting it with anything real, and operations for running one
after that.
Failure modes
A trimmed tool result leaves a silent gap
How to notice it
The final answer is missing a fact an earlier tool call actually returned, with no error and no retry: the context policy dropped it before a later call, and nothing downstream says so.
How to test for it
Run the same scripted model through a generous context policy and a tight one on the same question and compare the final text word for word; this page's own tests do exactly this.
A hook veto reads as the model refusing
How to notice it
A run stops short of an action, and it reads, from the transcript alone, like the model chose caution, when a hook actually blocked a call the model had already decided to make.
How to test for it
Read the trace, not the transcript. A veto is its own decided_by: "code" step; a model declining on its own is decided_by: "model". Confusing the two hides who is actually setting the policy.
A tool the model can see is one the registry will not run
How to notice it
The model calls a tool whose definition it was shown, and the call fails the way an unregistered name would, because the tool was advertised but never added to the allowlist that actually runs it.
How to test for it
Give the model a tool definition with no matching entry in the registry's allowlist and confirm the failure looks exactly like an unknown tool, not a special error: a mismatched allowlist should never be distinguishable from a typo.
Caps tuned for a different task cut every run short
How to notice it
Every run in a batch hits the step or token cap and returns a forced, partial answer, and the task looks fine in isolation: the caps were copied from a shorter task and never re-tuned.
How to test for it
Force a low cap on a task that genuinely needs more steps and confirm the forced answer is visibly marked, not indistinguishable from a real stop: this page's own tests do exactly this.
Logging the model's output is not logging the harness's decisions
How to notice it
A run goes wrong and the only record is what the model said (not which cap fired, what a policy trimmed, or which hook denied a call) so nobody can tell whether the model or the harness caused it.
How to test for it
Read what actually gets logged for one run end to end and check whether a cap, a trim, or a veto shows up in it at all, or only the text the model produced.
Cost and latency
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
2Model calls, best case (act, tool, stop)
5Model calls, worst case (step cap reached)
~180–310Tokens in, one action round
~0.6sWall time, one round trip
Compared with single agent (level 5, the same tools)The model-call shape is identical to single agent; a harness configuration changes how much of the growing context each call actually sees, not how many calls happen. A tight context policy can cost fewer tokens per call at the same step count, at the cost of what the model can still recall.
agent_harness answers a question about the documents and cites what it used, the same task
single_agent is scored on, so scripts/eval_run.py scores it against the site’s own
60-question set: run python scripts/eval_run.py --example agent_harness --model <spec> --dry
to project the cost of a full run before spending anything. What is specific to this technique is
running the same 60 questions through two harness configurations and comparing the two result
files directly: a gap in citation hit rate or score between a generous context policy and a tight
one is not model variance, it is what the harness cost the run.
Run it
What to monitor
Which cap ends a run (step, token, or a real stop), how often a hook vetoes a call the model chose, and how much of the growing context a policy actually keeps versus trims on a typical run: numbers a dashboard showing only the final answer will never surface.
Cost at volume
Tracks the same thing single agent's does (how many actions a question actually needs) plus one more: a tighter context policy costs fewer tokens per call at any given step count, so two harnesses running the identical loop can differ in spend without differing in step count at all.
How it fails in production
A context policy trims a fact a later step still needed, and the run finishes anyway with a plausible but wrong answer; or a hook denies a call quietly enough that a person reading only the final text never learns anything was blocked.
What to log
Which cap fired if any, every hook decision and its reason, what a context policy actually dropped on each call, and the tool registry's allowlist at the time of the run: a harness that only logs the model's own output cannot be debugged when the harness itself is what went wrong.
Try it
Use it
Pick two agent products (a coding agent and a research or deep-research tool, say) and try to name, for each, its step or turn cap, whether it asks before an irreversible action, and what it says when it hits a limit. That is the harness, not the model, and most products document at least part of it.
Build it
Run python -m examples.agent_harness --model stub:scripted from the repo root. The same model searches, looks up the part, and stops with a priced, warranty-scoped answer, all of it the harness's doing rather than the model's. Then open tests/test_example_agent_harness.py and change trim_to_budget(15) to a much larger number in the harness-changes-the-outcome test. Does the answer come back the same as the generous policy's?
Either lane
Pick one of the failure modes above and try to cause it on purpose, using the pattern in tests/test_example_agent_harness.py.
Build it
Read READ_ONLY_HEADERS and is_read_only in examples/common/bench.py, then SafetyEnvelope and Approval right after them. That is the same allowlist-plus-hook shape this page's own ToolRegistry and Hook build, in a setting where the one class of command that actually energizes a board needs a person's approval naming the set point before it runs at all.