# The agent harness

_Level 05 · Agent loops · sourced_

Everything around the model in an agent: the loop, tools, context handling, permissions, caps and sandbox.

## Conceptual architecture: The harness makes the loop executable.

Every model request passes through application controls before it affects the world.

- **Goal + context:** Task, instructions, selected history
- **Model decision:** Request a tool or return an answer
- **Execution gate:** Arguments, permissions, budgets
- **Run allowed tool:** Bounded operation in the environment
- **Observe the result:** Return output or a useful error
- **Finish or hand back:** Return work, evidence, and gaps
- **Pause or refuse:** Approval needed, denied, or capped

Connections:
- Goal + context → context → Model decision
- Model decision → tool request → Execution gate
- Execution gate → allowed → Run allowed tool
- Run allowed tool → observation → Observe the result
- Observe the result → next decision → Model decision
- Model decision → final answer → Finish or hand back
- Execution gate → cannot proceed → Pause or refuse

Reasoning helps the model choose useful actions. The loop supplies feedback; the harness supplies execution, state, and enforced limits. None of those makes the answer automatically correct.
- **Control:** Tool output is evidence, not permission to take another action.
- **Stopping:** Finish, ask for help, or stop at a step, time, or cost limit.
- **Verification:** Inspect the environment and the final artifact, not just the model’s account of its work.

## Try this in a recipe
- [Investigate an incident with bounded tools](/gradient_ascent/recipes/incident-runbook.md): Let a model choose read-only diagnostic tools, then require an evidence-backed handoff within six calls.

## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a task through the runtime surrounding the model: context assembly, tool access, execution, and stopping. The same model can behave differently when these surrounding choices change.

**Assumptions:** Instructions describe expected behavior; runtime permissions and checks determine which actions can actually execute. Their configuration must match the intended task.

**Design choices:** Choose tools, persistence, limits, and review points for the consequences of the work. A read-only research assistant and a deployment agent need different boundaries.

**Request:** Help plan a weekend trip, but do not book or pay for anything.

**Starting evidence:** Inputs: budget $400, step-free access needed. Tools: read mock schedules and save itinerary drafts. Booking tools disabled.

**Action and control:** The harness selects relevant context, permits read/draft tools, limits searches, and records blocked actions around the model's choices.

**Stage records (authored, not executed):**

### Input record

Inputs: budget $400, step-free access needed. Tools: read mock schedules and save itinerary drafts. Booking tools disabled.

What changed: Establish the facts supplied for this version of the task.

### Design note

Choose tools, persistence, limits, and review points for the consequences of the work. A read-only research assistant and a deployment agent need different boundaries.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

The harness selects relevant context, permits read/draft tools, limits searches, and records blocked actions around the model's choices.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Draft itinerary with unresolved accessibility checks. No bookings. Search cap and missing evidence are reported in the handoff.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Inspect selected context, allowed tools, stop reason, and evidence that no booking or payment occurred.

If the result falls short:
When a tool fails, evidence is missing, or a limit is reached, preserve useful state and report the open issue. Ask for expanded authority only when the task actually requires it.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Reuse the pattern with your own sources and tools. Keep goal, context, execution authority, and verification distinct; the DUT-specific ban on shared-framework edits is one policy, not the definition of a harness.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Draft itinerary with unresolved accessibility checks. No bookings. Search cap and missing evidence are reported in the handoff.

**Change something — A hotel page instructs the agent to pay a deposit:** Treat page text as data; the disabled payment tool remains unavailable. Log the blocked proposal and ask the user about next steps.

**Decision:** Does a website instruction expand the agent's authority?

**Answer:** No; runtime permissions and user scope still govern.

**Why:** The harness is the execution environment and controls, not the itinerary instructions alone.

**Review criteria:** Inspect selected context, allowed tools, stop reason, and evidence that no booking or payment occurred.

**Recovery:** When a tool fails, evidence is missing, or a limit is reached, preserve useful state and report the open issue. Ask for expanded authority only when the task actually requires it.

**Adapt it:** Reuse the pattern with your own sources and tools. Keep goal, context, execution authority, and verification distinct; the DUT-specific ban on shared-framework edits is one policy, not the definition of a harness.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a task through the runtime surrounding the model: context assembly, tool access, execution, and stopping. The same model can behave differently when these surrounding choices change.

**Assumptions:** Instructions describe expected behavior; runtime permissions and checks determine which actions can actually execute. Their configuration must match the intended task.

**Design choices:** Choose tools, persistence, limits, and review points for the consequences of the work. A read-only research assistant and a deployment agent need different boundaries.

**Request:** Prepare a weekly portfolio report using approved project sources, and wait for review before distribution.

**Starting evidence:** Tools: read tracker, read prior reports, save draft. Source access is scoped. W12 tracker: Atlas delayed; Cedar has no fresh update. Recipients: project leads.

**Action and control:** The harness supplies current evidence, checks tool permissions, limits retries, records provenance, and pauses for version-specific approval.

**Stage records (authored, not executed):**

### Input record

Tools: read tracker, read prior reports, save draft. Source access is scoped. W12 tracker: Atlas delayed; Cedar has no fresh update. Recipients: project leads.

What changed: Establish the facts supplied for this version of the task.

### Design note

Choose tools, persistence, limits, and review points for the consequences of the work. A read-only research assistant and a deployment agent need different boundaries.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

The harness supplies current evidence, checks tool permissions, limits retries, records provenance, and pauses for version-specific approval.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Draft W12-v1: Atlas delayed; Cedar has no fresh update. Source gaps visible. Distribution blocked until this report and audience are approved.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Review source access, attempts, stop reason, draft provenance, approval scope, and simulated distribution record.

If the result falls short:
When a tool fails, evidence is missing, or a limit is reached, preserve useful state and report the open issue. Ask for expanded authority only when the task actually requires it.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Reuse the pattern with your own sources and tools. Keep goal, context, execution authority, and verification distinct; the DUT-specific ban on shared-framework edits is one policy, not the definition of a harness.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Draft W12-v1: Atlas delayed; Cedar has no fresh update. Source gaps visible. Distribution blocked until this report and audience are approved.

**Change something — A connector remains unavailable after the retry limit:** Stop retries, mark affected project status unverified, and hand over a partial draft. Do not invent data or bypass access policy.

**Decision:** Should the harness keep retrying until it can present a complete report?

**Answer:** No; honor the limit and expose missing evidence.

**Why:** Runtime limits, context policy, and approval enforcement determine behavior even with the same model and request.

**Review criteria:** Review source access, attempts, stop reason, draft provenance, approval scope, and simulated distribution record.

**Recovery:** When a tool fails, evidence is missing, or a limit is reached, preserve useful state and report the open issue. Ask for expanded authority only when the task actually requires it.

**Adapt it:** Reuse the pattern with your own sources and tools. Keep goal, context, execution authority, and verification distinct; the DUT-specific ban on shared-framework edits is one policy, not the definition of a harness.

An agent harness is everything around the model in an agent: the loop that calls it, the tool
definitions it is shown and the code that runs them, what goes into its next request, whether an
action needs approval, the caps on steps and tokens, the sandbox, and what gets logged. None of that
is the model. Anthropic defines an agent in one sentence: systems "where LLMs dynamically direct
their own processes and tool usage, maintaining control over how they accomplish tasks"[1],
and almost everything a builder builds sits outside it.

That is why the same model behaves very differently in a different harness: change the step cap, the
allowlist or the context policy and a run finishes, fails, or runs up a bill doing neither, with
nothing about the model different. Anthropic's Claude Code team calls this loop engineering and
defines a loop as "agents repeating cycles of work until a stop condition is met"[4].

Level 5 is where the harness first has real decisions to bound: the model decides both the action
and when to stop: see [single agent](/gradient_ascent/techniques/single-agent/) for the loop
itself.

The harness is also where several supporting topics meet:
[guardrails](/gradient_ascent/techniques/guardrails/) check inputs, outputs, and proposed actions;
[human approval](/gradient_ascent/techniques/human-in-the-loop/) handles actions that need a person's decision;
[context engineering](/gradient_ascent/techniques/context-engineering/) shapes the next request;
and [observability](/gradient_ascent/techniques/observability/) records what happened.
[Evaluation](/gradient_ascent/techniques/evals/) checks the resulting system, while
[cost controls](/gradient_ascent/techniques/cost-optimization/) bound its work.
These are design choices around the loop, not capabilities guaranteed by the word “harness.”
Guardrail checks complement permission boundaries and sandboxing; they do not replace them.

This page is sourced, not measured: what the harnesses below do comes from their makers' own
documentation, and no run under one has been recorded and scored here. It is illustrated.

## Worked example: a test automation framework

A team uses a shared Python test automation framework. Each project represents a device under
test (DUT), follows the same project-file conventions, and uses the same or similar instruments.
The framework provides measurement methods, unit conversions, CSV export, and instrument drivers.
The drivers use SCPI, but project code treats instruments as black boxes through the framework's
interfaces. Python files contain project code; YAML/JSON files hold configuration such as instrument
settings, test parameters, limits, and sequences, according to the framework's conventions.

**Claude Code supplies the agent harness; the shared framework supplies the domain interfaces
and conventions.** Claude Code provides the model's read, edit, and command-execution loop.[8]
The agent uses that loop to create and refine a DUT project. The user reviews the files, performs
hardware testing, and returns logs and observations. The agent does not operate the instruments.

This is a specified workflow, not a recorded implementation or hardware validation result.
Simulation is not an established capability of this framework. Mocked function outputs could be
considered later, but are not assumed here.

The walkthrough below illustrates the agent workflow with scripted responses and sample artifacts.
It does not simulate instrument physics, execute project code, or call a model. Use **Watch it**
to follow the task, **Change something** to explore a missing requirement or new helper, and
**Try a decision** to check your understanding. The detailed reference follows the walkthrough.

**Overview:** Imagine your team already has a Python test framework and several past device projects. You need a project for a new device under test (DUT), with different requirements but familiar instruments and conventions. This walkthrough follows a coding agent from reading that context to handing over generated files for a person to review and test.

**Task:** Use reusable instructions in CLAUDE.md, the new DUT brief, framework documentation, and a suitable reference project to prepare Python tests, configuration, and Markdown documentation.

**What to look for:** Watch the plan become a scoped implementation, see a missing requirement or proposed helper trigger a decision, and distinguish generated files from evidence that the real measurements work.

**Adapt it:** The reusable pattern is context → proposed work → execution within authority → evidence → handoff. In this team, shared-framework edits, new helpers, and instrument access require separate authorization. Another project can preauthorize routine edits or safe checks. Choose boundaries around ownership, reversibility, and consequences rather than copying every restriction.

**Guided walkthrough:** Follow the DUT project through context, plan, approval, generation, non-hardware checks, and human handoff. Change a missing requirement or new-helper condition, then decide whether a new helper needs separate approval. Responses and check results are scripted illustrations, not model calls or executed validation.

### Go deeper: instructions, approval boundaries, code, and evaluation

### Start with an ordinary request

The user supplies two starting documents. **`CLAUDE.md` holds reusable framework instructions;
`DUT_BRIEF.md` describes this particular DUT.** Establish the framework instructions once and
maintain them as conventions evolve. Write a new brief for each DUT.

| Document | Who provides it | What it contains |
| --- | --- | --- |
| `CLAUDE.md` | User or framework maintainer; reused across DUT projects | Framework reference paths, project conventions, reuse rules, approval boundaries, permitted checks, and required deliverables. |
| `DUT_BRIEF.md` | User; specific to the new DUT | Required tests, differences from past DUTs, candidate reference projects, known instruments and configuration, and questions still to resolve. |
| `PROJECT_PLAN.md` | Agent drafts; user approves before implementation | Selected reference or template, proposed files and changes, framework tools to reuse, non-hardware checks, and approval requests. |
| `PROJECT_STATUS.md` | Agent creates and maintains from actual work and user feedback | Files created, capabilities, checks performed, hardware-validation status, limitations, and unresolved issues. |

The agent also needs access to framework documentation, framework source, previous projects,
and the standard template. These remain the reference material; the two starting files do not
replace them. Point to their actual locations in `CLAUDE.md`, and ask the agent to read
`DUT_BRIEF.md` when starting the task. The brief, plan, and status filenames are conventions for
this example, not special files automatically understood by every agent.

Configure permissions separately. Instructions in Markdown describe the boundaries; they do
not themselves make the framework read-only or prevent instrument access.

With those files in place, the user can give this request:

“Read CLAUDE.md and DUT_BRIEF.md, then create a project for this new DUT using our shared Python test framework. Use the closest past
project where one is suitable; otherwise use the standard template. Here is my description of how
this DUT differs. Read the framework documentation and reuse its existing tools and instrument
interfaces. Ask me about missing requirements before proceeding. Show me your proposed files and
changes in PROJECT_PLAN.md, and wait for my approval before generating the project. Produce Python code, YAML/JSON
configuration, and Markdown documentation, including PROJECT_STATUS.md, for my review. Obtain separate approval before creating any new project-local tool. Do not modify the shared framework or
connect to instruments.”

### What belongs where

| Part | Role in this example |
| --- | --- |
| Claude | Interprets the DUT differences, asks questions, proposes a plan, and drafts revisions. |
| Claude Code | Provides context management, file editing, command execution, and permission controls around the model. |
| Past projects and documentation | Supply the closest starting point, APIs, file conventions, and established patterns. A template is the fallback. |
| Shared Python framework | Provides reusable tools, measurement methods, unit conversion, CSV export, and instrument interfaces backed by SCPI drivers. |
| New DUT project | Contains the generated Python files, YAML/JSON configuration, and Markdown documentation. |
| Non-hardware checks | Check Python syntax and YAML/JSON validity without connecting to instruments. |
| User | Approves the plan, reviews the files, tests with real instruments, and supplies feedback. |

The agent can generate code that imports existing framework functions. Each function does not
need its own model-tool definition. The project's use of an instrument API does not authorize
the agent to execute it against connected equipment.

### The feedback loop in practice

1. **Understand:** read the framework documentation, candidate reference projects, and the written
   description of this DUT's differences. Ask about missing requirements before filling them in.
2. **Plan and wait:** identify the closest suitable project or the standard template. Propose the
   files, intended changes, framework tools to reuse, and non-hardware checks. Obtain the user's
   approval before generating project files and code.
3. **Generate:** create the approved Python files, YAML/JSON configuration, and Markdown documents
   in the new project's folder. Reuse the framework's existing tools wherever applicable.
4. **Check without hardware:** run approved syntax and configuration checks that cannot connect
   to instruments. Do not import or execute project setup code unless its lack of hardware access
   is established. Report the checks performed and their results; do not label the project as
   hardware-tested.
5. **Hand over:** the user reviews the project, runs it on real instruments, and supplies logs,
   results, and observations. The documentation distinguishes generated work from verified behavior.
6. **Refine:** use that feedback to revise the project within the approved scope. Ask about newly
   missing requirements, and seek approval for changes that cross the boundaries below.

The agent automates project creation and revision. Hardware testing remains a human-controlled
step. A passing syntax or configuration check does not establish that a measurement is correct.

### Approval boundaries and hard controls

**The shared framework must not be modified without explicit authorization.** If a new framework
feature is absolutely required, the agent must explain the requirement, why existing capabilities
cannot satisfy it, and the proposed change, then wait for the user's decision. Plan approval for a
DUT project is not blanket permission to modify the framework.

**A new project-local tool also requires approval.** Before creating one, explain the gap, which
framework tools were considered, and why a new tool is necessary. Placing a helper in the project
folder does not bypass this rule. Missing requirements must be asked about first, rather than
silently guessed or left as unapproved TODOs.

These are required boundaries, but a written instruction alone is not a hard enforcement mechanism.
Read-only access to the shared framework is a proposed control whose availability still needs to
be confirmed. The intended setup gives the agent write access only to the approved DUT project,
withholds live instrument access, and keeps permission controls outside files it can rewrite.
Any authorized framework change would need a separately scoped exception. This page does not claim
those controls are already implemented.

### Markdown documentation to hand over

- **Capabilities and created files:** supported tests, what was generated, and where each part lives.
- **Differences from the reference:** what changed for this DUT and why.
- **Configuration and operation:** parameters, limits, expected instruments, connections, setup,
  cleanup, and instructions for the user to run the project through the framework.
- **Requirements checklist:** each requested test mapped to its code, configuration, and validation status.
- **Validation record:** checks the agent actually ran, followed by hardware results the user supplies.
- **Open questions, limitations, and approvals:** unresolved issues and any requested tool or framework changes.

### How the concepts fit together

The harness coordinates these concepts during one task: **turn a DUT brief into a reviewable
project using an existing framework.** Some parts describe what the model sees, some decide
what may happen, and others establish what actually happened. The controls below describe the
intended setup; they are not a claim that custom checks or restrictions already exist.

| Concept | Where it appears in this DUT workflow | What it contributes |
| --- | --- | --- |
| [Instructions and prompting](/gradient_ascent/techniques/prompt-engineering/) | `CLAUDE.md` sets reusable rules; the request and `DUT_BRIEF.md` define the task. | Tell the agent to reuse framework functions, ask about unknowns, and deliver reviewable files. Instructions express policy; they do not enforce permissions. |
| [Context engineering](/gradient_ascent/techniques/context-engineering/) | Select the relevant API documentation, closest past project, DUT differences, approved plan, and latest feedback for the next model call. | Keep current requirements and approvals available as the conversation grows. Old project values are reference material, not automatically valid limits for this DUT. |
| [The agent loop](/gradient_ascent/techniques/single-agent/) | Read, propose, wait for approval, edit, check, inspect results, and revise. | The model chooses its next action within the allowed scope; the harness executes permitted actions and returns their results. Pause for missing requirements or an approval decision. |
| [Tools](/gradient_ascent/techniques/function-calling/) and [code execution](/gradient_ascent/techniques/code-execution/) | File reads, edits, and approved non-hardware checks are actions available to the coding agent. Generated Python calls the framework APIs later when the user runs it. | Distinguish an agent tool from a Python function used by the resulting project. Writing an instrument call does not grant permission to execute it. |
| [Guardrails](/gradient_ascent/techniques/guardrails/) | Proposed checks inspect intended actions and generated files for disallowed paths, direct instrument access, missing required configuration, or unapproved helpers. | Reject a prohibited action or flag work for correction. Syntax and schema checks can be deterministic; judging whether an existing tool meets a requirement may still need human review. These checks must be implemented and tested. |
| Permissions and sandboxing | Configure project-only writes, protected framework files, and an environment without live instrument access. | Bound what executed code can actually touch, even if the model proposes otherwise. Read-only framework access remains to be confirmed; a path check alone is not a complete sandbox. |
| [Human approval](/gradient_ascent/techniques/human-in-the-loop/) | The user approves `PROJECT_PLAN.md`, separately decides on any new helper or framework change, and controls hardware testing. | Resolve a decision the agent cannot authorize for itself. Approval is scoped to the stated change; a declined request leaves the boundary in place. |
| [Observability](/gradient_ascent/techniques/observability/) | Retain file diffs, commands, check results, approval decisions, and reasons for blocked actions; summarize progress in `PROJECT_STATUS.md`. | Explain why the run changed a file, stopped, or failed. The status document is a readable summary, not a substitute for the underlying execution record. |
| [Evaluation](/gradient_ascent/techniques/evals/) | Independently compare the output with the DUT requirements, framework conventions, approval record, and reported validation status. | Assess the agent's work. User-run hardware tests assess the resulting measurement behavior; passing syntax checks establishes neither of these on its own. |
| [Cost and stop controls](/gradient_ascent/techniques/cost-optimization/) | Set supported step, time, or spend limits, plus a policy for repeated failed checks and unresolved requirements. | Bound revision work and hand back a partial result with a clear reason for stopping. Exact budgets have not been selected for this example. |
| [Persistent task state](/gradient_ascent/techniques/memory/) | Save the approved plan, project status, and user feedback; explicitly load them when resuming. | Carry decisions between sessions without assuming the model remembers them. Saved files only help when their relevant contents reach the next request. |

### One proposed helper, several different controls

Suppose the agent believes it needs a new unit-conversion helper for the DUT:

1. **Context and tools:** it reads the framework's existing conversion API and relevant past code.
   If those already meet the requirement, it uses them in the generated project.
2. **Guardrail and approval:** if it proposes a new helper, a configured policy check should pause
   that creation until a specific approval exists. The agent explains the gap and asks the user.
   Without an implemented check, this remains an instruction the agent is expected to follow.
3. **Permissions:** approval for a helper in the project does not unlock the shared framework or
   instrument access. A framework change would require its own authorization and scoped access.
4. **Execution and observation:** after approval, it writes the helper within scope, runs only
   approved non-hardware checks, and records the change, approval, and actual results.
5. **Evaluation and feedback:** the reviewer checks the conversion against the agreed requirement.
   The user performs any required hardware validation and returns findings. The agent revises
   within scope or stops and reports the next decision it needs.

**The guardrail checks the proposal; the person authorizes an exception; permissions constrain
execution; logs record the outcome; evaluation judges whether the result meets the requirement.**
The harness brings these together around the model's repeated calls.

### Relating this to the code and run below

The runnable demonstration below uses a document lookup task, not this DUT framework. Its
`ContextPolicy` corresponds to selecting the documentation and decisions the model sees;
`ToolRegistry` corresponds to the actions the coding agent can call; a `Hook` illustrates checking
a proposed action before execution; and the caps bound the loop. A veto hook alone is not a human
approval workflow: that also needs a pause, a recorded decision, and a way to resume within scope.
The trace illustrates observability. The DUT diagram describes how these responsibilities would
apply to your project workflow, without claiming the demonstration implements its controls.

You do not need every technique on the site to start this workflow. Reading repository files does
not by itself establish a RAG system, a Markdown instruction file is not automatically a packaged
skill, and using Python APIs does not require MCP. Those are separate choices if retrieval,
reusable procedures, or external tool connections become necessary.

_The web page for this technique includes an interactive step-through of Level 5 · The agent harness. The same steps are described in the sections below._

## Practical guidance

If you use a chat app and will never run an agent, skip this page. A harness is the code wrapped
around the model, written by whoever built the agent product, and there is no box for you to type
in. The pages that are yours are [coding agents](/gradient_ascent/techniques/coding-agents/) and
[always-on assistants](/gradient_ascent/techniques/agent-teammates/).

If you do operate an agent product, one thing here earns your time: when an agent behaves badly,
the fix is usually a setting rather than a better prompt. Four settings, and what each one looks
like when it is the cause.

**What it remembers.** Anthropic's Claude Agent SDK documents automatic compaction firing as a
long session grows: "When the context window approaches its limit, the SDK automatically compacts
the conversation: it summarizes older history to free space, keeping your most recent exchanges
and key decisions intact"[2]. Anthropic's engineering blog describes tool result
clearing, a lighter version of the same idea, as one of the "safest lightest touch forms of
compaction"[3]. An agent that dropped your constraint two hours into a session did not
ignore it; it summarized it away. Restate the constraint in your next message rather than starting
the whole task again.

**What it may run without asking.** OpenAI's Codex documentation states the split: "Sandboxing and
approvals are different controls that work together. The sandbox defines technical boundaries. The
approval policy decides when the agent must stop and ask before crossing them"[6],
enforced on macOS "using the built-in Seatbelt framework"[6]. Between asking every time
and never asking, a policy can "keep specific approval prompt categories interactive while
automatically rejecting others"[7]. Start on the setting that asks, and loosen one
category at a time once you have watched what that category actually does.

**What it costs before it stops.** A cap on steps or spend is a harness decision, not a model one,
and a run that hits one usually ends mid-task with no error: see
[single agent](/gradient_ascent/techniques/single-agent/).

**What a session hands on.** A subagent may explore at length and return "only a condensed,
distilled summary of its work"[3] to the harness that spawned it. That is why an agent's
account of what it did can be thinner than what it did, and why
[memory](/gradient_ascent/techniques/memory/) is a separate setting from the summary.

Read those four in your own product's documentation before concluding a rough session was the
model's fault.

## Implementation details

The document lookup demonstration is the same loop [single agent](/gradient_ascent/techniques/single-agent/) runs —
act, check a cap, repeat: rebuilt so four moving parts are arguments to `run` instead of fixed in
the function body: a `ToolRegistry` (the definitions the model is shown, and the `allowed` set
checked before any of them runs: the split
[function calling](/gradient_ascent/techniques/function-calling/) makes for one call, made
reusable), a `ContextPolicy` (a function from the growing message history to whatever the next
request sends: [context engineering](/gradient_ascent/techniques/context-engineering/) applied
inside the loop rather than once before it), a `Hook` (a chance to veto a call the model already
chose, before the registry runs it: a silent version of what
[human approval](/gradient_ascent/techniques/human-in-the-loop/) does out loud), and the step and
token caps.

`trim_to_budget` is the context policy worth reading closely. It keeps every message except tool
results, and keeps only as many of the most recent tool results as fit under a token budget,
replacing older ones with a short placeholder rather than deleting them silently: a small version
of what Anthropic calls tool result clearing[3]:

`examples/agent_harness/run.py` (lines 75-112)

```python
def trim_to_budget(budget_tokens: int) -> ContextPolicy:
    """The tight policy: keeps every non-tool-result message, and as many of the most recent
    tool results as fit under `budget_tokens`, dropping older ones first. Real context policies
    trim the same way -- see this page's Use it lane for how Anthropic describes tool result
    clearing and compaction -- this one trims by a plain token count to keep the point readable
    in a few lines.

    A tool result is recognized by its role, `tool`, never by how its text opens: the question is
    a user message, so a policy that matched on text could drop a question that happened to begin
    "Result of ...", leaving the model answering something it can no longer see. A trimmed result
    keeps its call id, because every tool call in the history still needs an answer."""

    def policy(messages: list[Message]) -> list[Message]:
        result_idx = [i for i, m in enumerate(messages) if _is_tool_result(m)]
        kept: set[int] = set()
        used = 0
        for i in reversed(result_idx):
            cost = count_tokens(content_text(messages[i].content))
            if used + cost > budget_tokens:
                break
            used += cost
            kept.add(i)
        out = []
        for i, m in enumerate(messages):
            if i in result_idx and i not in kept:
                out.append(
                    Message(
                        role="tool",
                        content="[earlier tool result trimmed by the context policy]",
                        tool_call_id=m.tool_call_id,
                        tool_name=m.tool_name,
                    )
                )
            else:
                out.append(m)
        return out

    return policy
```

It runs fresh on every model call, not once at the start, which is what lets the two runs below
diverge partway through instead of only at the first prompt. It also leaves the opening request
alone: it recognizes a tool result by how the text opens, so without that guard a question starting
the same way would be trimmed and the model would be answering something it could no longer see.
`deny_after` is the hook worth reading next, three lines, and the point is that it runs after the
model has already decided:

`examples/agent_harness/run.py` (lines 129-142)

```python
def deny_after(allowed_calls: int) -> Hook:
    """A hook for demonstration and testing: allows the first `allowed_calls` tool calls the
    model attempts, vetoes every one after. A real hook would read the call's own name and
    arguments; this one only counts, to keep the point -- a hook can block an action the model
    already decided to take -- in three lines."""
    seen = {"n": 0}

    def hook(call: ToolCall) -> tuple[bool, str]:
        seen["n"] += 1
        if seen["n"] > allowed_calls:
            return False, f"tool budget of {allowed_calls} call(s) already spent"
        return True, ""

    return hook
```

A real hook would look at the call's own name and arguments instead of just counting; running one
inside a sandbox that actually isolates what a tool may touch, rather than a check like this one,
is [code execution](/gradient_ascent/techniques/code-execution/)'s territory, and what gets loaded
into the model's instructions in the first place (which skill, not just which tool) is
[skills](/gradient_ascent/techniques/skills/)'.

`run` is the loop these parts plug into: ask the model through whatever the context policy
currently allows it to see, and if it calls tools, check each one against the hook and the
registry's allowlist before running it, record what happened, and go around again until the model
stops or a cap does.

`examples/agent_harness/run.py` (lines 145-197)

```python
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir=DEFAULT_CORPUS_DIR,
    registry: ToolRegistry = DEFAULT_REGISTRY,
    context_policy: ContextPolicy = keep_everything,
    hook: Hook = allow_everything,
    max_steps: int = MAX_STEPS,
    max_tokens: int = MAX_TOKENS,
) -> Answer:
    del embedder  # this harness retrieves through its tools, not a vector index
    sections = load_sections(corpus_dir)
    messages = [Message(role="system", content=SYSTEM), Message(role="user", content=question)]
    citations: list[str] = []
    tokens_used = 0

    for _ in range(max_steps):
        completion = model.complete(context_policy(messages), tools=registry.definitions, max_tokens=400)
        tokens_used += completion.tokens_in + completion.tokens_out

        if not completion.tool_calls:
            record_completion(tracer, decided_by="model", title="Model stops and answers", completion=completion)
            return Answer.from_text(completion.text, retrieved_sources=citations)

        calls_desc = ", ".join(f"{c.name}({c.arguments})" for c in completion.tool_calls)
        record_completion(tracer, decided_by="model", title="Model picks an action", completion=completion, detail=calls_desc)
        turn, calls = assistant_turn(completion, len(messages))
        messages.append(turn)

        for call in calls:
            allowed, reason = hook(call)
            if not allowed:
                tracer.record(kind="code", decided_by="code", title="Hook vetoes the call", detail=reason)
                messages.append(tool_result(call, f"Denied: {reason}"))
                continue
            if call.name not in registry.allowed:
                result_text, cites = toolkit.unknown_tool(call.name)
            else:
                result_text, cites = registry.call(call, sections)
            citations.extend(cites)
            tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])
            messages.append(tool_result(call, result_text))

        if tokens_used >= max_tokens:
            reason = f"token budget reached: {tokens_used} >= {max_tokens}"
            final = force_final(context_policy(messages), model, tracer, reason=reason, max_tokens=400)
            return Answer.from_text(final.text, retrieved_sources=citations)

    final = force_final(context_policy(messages), model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
    return Answer.from_text(final.text, retrieved_sources=citations)
```

Every tool call, its arguments, and the decision to stop are `decided_by: "model"`; running a tool,
a hook's veto, and forcing a final answer when a cap is reached are always `decided_by: "code"`:
the same split `single_agent`'s example makes, with two more kinds of code-decided step than that
one has.

`tests/test_example_agent_harness.py` runs the same scripted model twice with only the context
policy changed, and the two runs answer differently: confidently citing the warranty term under a
generous policy, saying it could not confirm the term under a tight one. Nothing about the model's
own logic changed between the two runs; only what the harness let it see did. Read that for what it
is: a scripted stand-in, written to answer from whatever the harness left in front of it, so what
the test proves is the mechanism, not a measurement of how much a real model's answers move. The
size of that effect is what the eval below is for, and no run of it exists yet. The same file scripts a hook that vetoes a call the
model already committed to, and checks the veto shows up in the trace as the harness's own
decision, and a tool the registry advertises but will not run, and checks it fails exactly the way
an unknown tool does.

Run it yourself:

`examples/agent_harness/README.md` (lines 15-15)

```text
python -m examples.agent_harness --model stub:scripted
```

## When you do not need this

Try [single agent](/gradient_ascent/techniques/single-agent/) first if you have not seen the basic
loop yet: this page assumes you have, and is about what surrounds it, not the loop itself.

Move to thinking about the harness deliberately once an agent runs past a one-off demo: choosing
the caps, the allowlist, the context policy and the approval settings on purpose is what turns a
loop that happens to work into a system somebody can operate and debug: see
[safety, privacy and governance](/gradient_ascent/techniques/safety/) for testing one before
trusting it with anything real, and [operations](/gradient_ascent/techniques/ops/) for running one
after that.

## Failure modes

### A trimmed tool result leaves a silent gap

- **How to notice it:** The final answer is missing a fact an earlier tool call actually returned, with no error and no retry: the context policy dropped it before a later call, and nothing downstream says so.
- **How to test for it:** Run the same scripted model through a generous context policy and a tight one on the same question and compare the final text word for word; this page's own tests do exactly this.

### A hook veto reads as the model refusing

- **How to notice it:** A run stops short of an action, and it reads, from the transcript alone, like the model chose caution, when a hook actually blocked a call the model had already decided to make.
- **How to test for it:** Read the trace, not the transcript. A veto is its own decided_by: "code" step; a model declining on its own is decided_by: "model". Confusing the two hides who is actually setting the policy.

### A tool the model can see is one the registry will not run

- **How to notice it:** The model calls a tool whose definition it was shown, and the call fails the way an unregistered name would, because the tool was advertised but never added to the allowlist that actually runs it.
- **How to test for it:** Give the model a tool definition with no matching entry in the registry's allowlist and confirm the failure looks exactly like an unknown tool, not a special error: a mismatched allowlist should never be distinguishable from a typo.

### Caps tuned for a different task cut every run short

- **How to notice it:** Every run in a batch hits the step or token cap and returns a forced, partial answer, and the task looks fine in isolation: the caps were copied from a shorter task and never re-tuned.
- **How to test for it:** Force a low cap on a task that genuinely needs more steps and confirm the forced answer is visibly marked, not indistinguishable from a real stop: this page's own tests do exactly this.

### Logging the model's output is not logging the harness's decisions

- **How to notice it:** A run goes wrong and the only record is what the model said (not which cap fired, what a policy trimmed, or which hook denied a call) so nobody can tell whether the model or the harness caused it.
- **How to test for it:** Read what actually gets logged for one run end to end and check whether a cap, a trim, or a veto shows up in it at all, or only the text the model produced.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, best case (act, tool, stop):** 2
- **Model calls, worst case (step cap reached):** 5
- **Tokens in, one action round:** ~180–310
- **Wall time, one round trip:** ~0.6s

**Compared with single agent (level 5, the same tools).** The model-call shape is identical to single agent; a harness configuration changes how much of the growing context each call actually sees, not how many calls happen. A tight context policy can cost fewer tokens per call at the same step count, at the cost of what the model can still recall.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

`agent_harness` answers a question about the documents and cites what it used, the same task
`single_agent` is scored on, so `scripts/eval_run.py` scores it against the site's own
60-question set: run `python scripts/eval_run.py --example agent_harness --model <spec> --dry`
to project the cost of a full run before spending anything. What is specific to this technique is
running the same 60 questions through two harness configurations and comparing the two result
files directly: a gap in citation hit rate or score between a generous context policy and a tight
one is not model variance, it is what the harness cost the run.

## Run it

**What to monitor.** Which cap ends a run (step, token, or a real stop), how often a hook vetoes a call the model chose, and how much of the growing context a policy actually keeps versus trims on a typical run: numbers a dashboard showing only the final answer will never surface.

**Cost at volume.** Tracks the same thing single agent's does (how many actions a question actually needs) plus one more: a tighter context policy costs fewer tokens per call at any given step count, so two harnesses running the identical loop can differ in spend without differing in step count at all.

**How it fails in production.** A context policy trims a fact a later step still needed, and the run finishes anyway with a plausible but wrong answer; or a hook denies a call quietly enough that a person reading only the final text never learns anything was blocked.

**What to log.** Which cap fired if any, every hook decision and its reason, what a context policy actually dropped on each call, and the tool registry's allowlist at the time of the run: a harness that only logs the model's own output cannot be debugged when the harness itself is what went wrong.

## Try it

1. **Use it.** Pick two agent products (a coding agent and a research or deep-research tool, say) and try to name, for each, its step or turn cap, whether it asks before an irreversible action, and what it says when it hits a limit. That is the harness, not the model, and most products document at least part of it.
2. **Build it.** Run python -m examples.agent_harness --model stub:scripted from the repo root. The same model searches, looks up the part, and stops with a priced, warranty-scoped answer, all of it the harness's doing rather than the model's. Then open tests/test_example_agent_harness.py and change trim_to_budget(15) to a much larger number in the harness-changes-the-outcome test. Does the answer come back the same as the generous policy's?
3. **Either lane.** Pick one of the failure modes above and try to cause it on purpose, using the pattern in tests/test_example_agent_harness.py.
4. **Build it.** Read READ_ONLY_HEADERS and is_read_only in examples/common/bench.py, then SafetyEnvelope and Approval right after them. That is the same allowlist-plus-hook shape this page's own ToolRegistry and Hook build, in a setting where the one class of command that actually energizes a board needs a person's approval naming the set point before it runs at all.


## Sources

1. [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic, 2024-12-19 (accessed 2026-09-19)
2. [How the agent loop works](https://code.claude.com/docs/en/agent-sdk/agent-loop) — Anthropic (Claude Agent SDK documentation) (accessed 2026-09-19)
3. [Effective context engineering for AI agents](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) — Anthropic (Engineering blog) (accessed 2026-09-19)
4. [Loop engineering: getting started with loops](https://claude.com/blog/getting-started-with-loops) — Anthropic (Claude blog) (accessed 2026-09-19)
5. [Running agents](https://openai.github.io/openai-agents-python/running_agents/) — OpenAI (Agents SDK documentation) (accessed 2026-09-19)
6. [Sandboxing](https://learn.chatgpt.com/docs/sandboxing) — OpenAI (Codex documentation) (accessed 2026-09-19)
7. [Agent approvals & security](https://learn.chatgpt.com/docs/agent-approvals-security) — OpenAI (Codex documentation) (accessed 2026-09-19)
8. [How Claude Code works](https://code.claude.com/docs/en/how-claude-code-works) — Anthropic (Claude Code documentation) (accessed 2026-09-19)


Last reviewed 2026-09-19.
