Level 05 · Agent loops

The agent harness

Everything around the model in an agent: the loop, tools, context handling, permissions, caps and sandbox.

Sourced

How it works · conceptual architecture

The harness makes the loop executable.

Every model request passes through application controls before it affects the world.

Step / conditionInformation / relationshipReturn / repeatHighlighted box: model
The harness makes the loop executable.Goal + context → context → Model decision. Model decision → tool request → Execution gate. Execution gate → allowed → Run allowed tool. Run allowed tool → observation → Observe the result. Observe the result → next decision → Model decision. Model decision → final answer → Finish or hand back. Execution gate → cannot proceed → Pause or refuse.contexttool requestallowedobservationnext decisionfinal answercannot proceedAGoal + contextTask, instructions, selectedhistoryBModel decisionRequest a tool or return ananswerCExecution gateArguments, permissions,budgetsDRun allowed toolBounded operation in theenvironmentEObserve the resultReturn output or a usefulerrorFFinish or hand backReturn work, evidence, andgapsGPause or refuseApproval needed, denied, orcapped
A
Goal + context

Task, instructions, selected history

  • context B · Model decision
B
Model decision

Request a tool or return an answer

  • tool request C · Execution gate
  • final answer F · Finish or hand back
C
Execution gate

Arguments, permissions, budgets

  • allowed D · Run allowed tool
  • cannot proceed G · Pause or refuse
D
Run allowed tool

Bounded operation in the environment

  • observation E · Observe the result
E
Observe the result

Return output or a useful error

  • next decision B · Model decision
F
Finish or hand back

Return work, evidence, and gaps

    G
    Pause or refuse

    Approval needed, denied, or capped

      Reasoning helps the model choose useful actions. The loop supplies feedback; the harness supplies execution, state, and enforced limits. None of those makes the answer automatically correct.
      The details that change the design

      Control

      Tool output is evidence, not permission to take another action.

      Stopping

      Finish, ask for help, or stop at a step, time, or cost limit.

      Verification

      Inspect the environment and the final artifact, not just the model’s account of its work.

      Apply this to your project

      Describe your task to your own model and use The agent harness as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

      Go deeper: practical guidance, failure modes, and implementation

      An agent harness is everything around the model in an agent: the loop that calls it, the tool definitions it is shown and the code that runs them, what goes into its next request, whether an action needs approval, the caps on steps and tokens, the sandbox, and what gets logged. None of that is the model. Anthropic defines an agent in one sentence: systems “where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks”[1], and almost everything a builder builds sits outside it.

      That is why the same model behaves very differently in a different harness: change the step cap, the allowlist or the context policy and a run finishes, fails, or runs up a bill doing neither, with nothing about the model different. Anthropic’s Claude Code team calls this loop engineering and defines a loop as “agents repeating cycles of work until a stop condition is met”[4].

      Level 5 is where the harness first has real decisions to bound: the model decides both the action and when to stop: see single agent for the loop itself.

      The harness is also where several supporting topics meet: guardrails check inputs, outputs, and proposed actions; human approval handles actions that need a person’s decision; context engineering shapes the next request; and observability records what happened. Evaluation checks the resulting system, while cost controls bound its work. These are design choices around the loop, not capabilities guaranteed by the word “harness.” Guardrail checks complement permission boundaries and sandboxing; they do not replace them.

      This page is sourced, not measured: what the harnesses below do comes from their makers’ own documentation, and no run under one has been recorded and scored here. It is illustrated.

      Worked example: a test automation framework

      A team uses a shared Python test automation framework. Each project represents a device under test (DUT), follows the same project-file conventions, and uses the same or similar instruments. The framework provides measurement methods, unit conversions, CSV export, and instrument drivers. The drivers use SCPI, but project code treats instruments as black boxes through the framework’s interfaces. Python files contain project code; YAML/JSON files hold configuration such as instrument settings, test parameters, limits, and sequences, according to the framework’s conventions.

      Claude Code supplies the agent harness; the shared framework supplies the domain interfaces and conventions. Claude Code provides the model’s read, edit, and command-execution loop.[8] The agent uses that loop to create and refine a DUT project. The user reviews the files, performs hardware testing, and returns logs and observations. The agent does not operate the instruments.

      This is a specified workflow, not a recorded implementation or hardware validation result. Simulation is not an established capability of this framework. Mocked function outputs could be considered later, but are not assumed here.

      The walkthrough below illustrates the agent workflow with scripted responses and sample artifacts. It does not simulate instrument physics, execute project code, or call a model. Use Watch it to follow the task, Change something to explore a missing requirement or new helper, and Try a decision to check your understanding. The detailed reference follows the walkthrough.

      CHOOSE YOUR PERSPECTIVE

      Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

      GUIDED WORKED EXAMPLE Scripted simulation · no model calls or hardware access

      From DUT brief to reviewable project.

      Follow one task. See what the agent proposes, what the harness controls, and where you decide.

      What you’ll walk through

      Imagine your team already has a Python test framework and several past device projects. You need a project for a new device under test (DUT), with different requirements but familiar instruments and conventions. This walkthrough follows a coding agent from reading that context to handing over generated files for a person to review and test.

      The task in this version

      Use reusable instructions in CLAUDE.md, the new DUT brief, framework documentation, and a suitable reference project to prepare Python tests, configuration, and Markdown documentation.

      What you’ll learn to check

      Watch the plan become a scoped implementation, see a missing requirement or proposed helper trigger a decision, and distinguish generated files from evidence that the real measurements work.

      The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

      How this applies beyond this test framework

      The reusable pattern is context → proposed work → execution within authority → evidence → handoff. In this team, shared-framework edits, new helpers, and instrument access require separate authorization. Another project can preauthorize routine edits or safe checks. Choose boundaries around ownership, reversibility, and consequences rather than copying every restriction.

      Engineering & technical workCreate a DUT project within the shared framework’s conventions and approval boundaries.
      Prefilled English request. You can edit it; this walkthrough always demonstrates the same scripted workflow.
      THE WORKFLOWReady when you are
      ContextInstructions + DUT brief + reference project
      Always in effectProject-only scopeFramework changes need authorizationNo live instrument accessIntended controls, illustrated here.
      THE VISIBLE WORKStart with a task
      01 → 06

      A project takes shape, one decision at a time.

      Send the request above. Then follow the agent’s work and make the approval decision yourself.

      Brief → Plan → Approval → Files → Evidence
      Go deeper: instructions, approval boundaries, code, and evaluation

      Start with an ordinary request

      The user supplies two starting documents. CLAUDE.md holds reusable framework instructions; DUT_BRIEF.md describes this particular DUT. Establish the framework instructions once and maintain them as conventions evolve. Write a new brief for each DUT.

      Document Who provides it What it contains
      CLAUDE.md User or framework maintainer; reused across DUT projects Framework reference paths, project conventions, reuse rules, approval boundaries, permitted checks, and required deliverables.
      DUT_BRIEF.md User; specific to the new DUT Required tests, differences from past DUTs, candidate reference projects, known instruments and configuration, and questions still to resolve.
      PROJECT_PLAN.md Agent drafts; user approves before implementation Selected reference or template, proposed files and changes, framework tools to reuse, non-hardware checks, and approval requests.
      PROJECT_STATUS.md Agent creates and maintains from actual work and user feedback Files created, capabilities, checks performed, hardware-validation status, limitations, and unresolved issues.

      The agent also needs access to framework documentation, framework source, previous projects, and the standard template. These remain the reference material; the two starting files do not replace them. Point to their actual locations in CLAUDE.md, and ask the agent to read DUT_BRIEF.md when starting the task. The brief, plan, and status filenames are conventions for this example, not special files automatically understood by every agent.

      Configure permissions separately. Instructions in Markdown describe the boundaries; they do not themselves make the framework read-only or prevent instrument access.

      With those files in place, the user can give this request:

      “Read CLAUDE.md and DUT_BRIEF.md, then create a project for this new DUT using our shared Python test framework. Use the closest past project where one is suitable; otherwise use the standard template. Here is my description of how this DUT differs. Read the framework documentation and reuse its existing tools and instrument interfaces. Ask me about missing requirements before proceeding. Show me your proposed files and changes in PROJECT_PLAN.md, and wait for my approval before generating the project. Produce Python code, YAML/JSON configuration, and Markdown documentation, including PROJECT_STATUS.md, for my review. Obtain separate approval before creating any new project-local tool. Do not modify the shared framework or connect to instruments.”

      What belongs where

      Part Role in this example
      Claude Interprets the DUT differences, asks questions, proposes a plan, and drafts revisions.
      Claude Code Provides context management, file editing, command execution, and permission controls around the model.
      Past projects and documentation Supply the closest starting point, APIs, file conventions, and established patterns. A template is the fallback.
      Shared Python framework Provides reusable tools, measurement methods, unit conversion, CSV export, and instrument interfaces backed by SCPI drivers.
      New DUT project Contains the generated Python files, YAML/JSON configuration, and Markdown documentation.
      Non-hardware checks Check Python syntax and YAML/JSON validity without connecting to instruments.
      User Approves the plan, reviews the files, tests with real instruments, and supplies feedback.

      The agent can generate code that imports existing framework functions. Each function does not need its own model-tool definition. The project’s use of an instrument API does not authorize the agent to execute it against connected equipment.

      The feedback loop in practice

      1. Understand: read the framework documentation, candidate reference projects, and the written description of this DUT’s differences. Ask about missing requirements before filling them in.
      2. Plan and wait: identify the closest suitable project or the standard template. Propose the files, intended changes, framework tools to reuse, and non-hardware checks. Obtain the user’s approval before generating project files and code.
      3. Generate: create the approved Python files, YAML/JSON configuration, and Markdown documents in the new project’s folder. Reuse the framework’s existing tools wherever applicable.
      4. Check without hardware: run approved syntax and configuration checks that cannot connect to instruments. Do not import or execute project setup code unless its lack of hardware access is established. Report the checks performed and their results; do not label the project as hardware-tested.
      5. Hand over: the user reviews the project, runs it on real instruments, and supplies logs, results, and observations. The documentation distinguishes generated work from verified behavior.
      6. Refine: use that feedback to revise the project within the approved scope. Ask about newly missing requirements, and seek approval for changes that cross the boundaries below.

      The agent automates project creation and revision. Hardware testing remains a human-controlled step. A passing syntax or configuration check does not establish that a measurement is correct.

      Approval boundaries and hard controls

      The shared framework must not be modified without explicit authorization. If a new framework feature is absolutely required, the agent must explain the requirement, why existing capabilities cannot satisfy it, and the proposed change, then wait for the user’s decision. Plan approval for a DUT project is not blanket permission to modify the framework.

      A new project-local tool also requires approval. Before creating one, explain the gap, which framework tools were considered, and why a new tool is necessary. Placing a helper in the project folder does not bypass this rule. Missing requirements must be asked about first, rather than silently guessed or left as unapproved TODOs.

      These are required boundaries, but a written instruction alone is not a hard enforcement mechanism. Read-only access to the shared framework is a proposed control whose availability still needs to be confirmed. The intended setup gives the agent write access only to the approved DUT project, withholds live instrument access, and keeps permission controls outside files it can rewrite. Any authorized framework change would need a separately scoped exception. This page does not claim those controls are already implemented.

      Markdown documentation to hand over

      • Capabilities and created files: supported tests, what was generated, and where each part lives.
      • Differences from the reference: what changed for this DUT and why.
      • Configuration and operation: parameters, limits, expected instruments, connections, setup, cleanup, and instructions for the user to run the project through the framework.
      • Requirements checklist: each requested test mapped to its code, configuration, and validation status.
      • Validation record: checks the agent actually ran, followed by hardware results the user supplies.
      • Open questions, limitations, and approvals: unresolved issues and any requested tool or framework changes.

      How the concepts fit together

      The harness coordinates these concepts during one task: turn a DUT brief into a reviewable project using an existing framework. Some parts describe what the model sees, some decide what may happen, and others establish what actually happened. The controls below describe the intended setup; they are not a claim that custom checks or restrictions already exist.

      Concept Where it appears in this DUT workflow What it contributes
      Instructions and prompting CLAUDE.md sets reusable rules; the request and DUT_BRIEF.md define the task. Tell the agent to reuse framework functions, ask about unknowns, and deliver reviewable files. Instructions express policy; they do not enforce permissions.
      Context engineering Select the relevant API documentation, closest past project, DUT differences, approved plan, and latest feedback for the next model call. Keep current requirements and approvals available as the conversation grows. Old project values are reference material, not automatically valid limits for this DUT.
      The agent loop Read, propose, wait for approval, edit, check, inspect results, and revise. The model chooses its next action within the allowed scope; the harness executes permitted actions and returns their results. Pause for missing requirements or an approval decision.
      Tools and code execution File reads, edits, and approved non-hardware checks are actions available to the coding agent. Generated Python calls the framework APIs later when the user runs it. Distinguish an agent tool from a Python function used by the resulting project. Writing an instrument call does not grant permission to execute it.
      Guardrails Proposed checks inspect intended actions and generated files for disallowed paths, direct instrument access, missing required configuration, or unapproved helpers. Reject a prohibited action or flag work for correction. Syntax and schema checks can be deterministic; judging whether an existing tool meets a requirement may still need human review. These checks must be implemented and tested.
      Permissions and sandboxing Configure project-only writes, protected framework files, and an environment without live instrument access. Bound what executed code can actually touch, even if the model proposes otherwise. Read-only framework access remains to be confirmed; a path check alone is not a complete sandbox.
      Human approval The user approves PROJECT_PLAN.md, separately decides on any new helper or framework change, and controls hardware testing. Resolve a decision the agent cannot authorize for itself. Approval is scoped to the stated change; a declined request leaves the boundary in place.
      Observability Retain file diffs, commands, check results, approval decisions, and reasons for blocked actions; summarize progress in PROJECT_STATUS.md. Explain why the run changed a file, stopped, or failed. The status document is a readable summary, not a substitute for the underlying execution record.
      Evaluation Independently compare the output with the DUT requirements, framework conventions, approval record, and reported validation status. Assess the agent’s work. User-run hardware tests assess the resulting measurement behavior; passing syntax checks establishes neither of these on its own.
      Cost and stop controls Set supported step, time, or spend limits, plus a policy for repeated failed checks and unresolved requirements. Bound revision work and hand back a partial result with a clear reason for stopping. Exact budgets have not been selected for this example.
      Persistent task state Save the approved plan, project status, and user feedback; explicitly load them when resuming. Carry decisions between sessions without assuming the model remembers them. Saved files only help when their relevant contents reach the next request.

      One proposed helper, several different controls

      Suppose the agent believes it needs a new unit-conversion helper for the DUT:

      1. Context and tools: it reads the framework’s existing conversion API and relevant past code. If those already meet the requirement, it uses them in the generated project.
      2. Guardrail and approval: if it proposes a new helper, a configured policy check should pause that creation until a specific approval exists. The agent explains the gap and asks the user. Without an implemented check, this remains an instruction the agent is expected to follow.
      3. Permissions: approval for a helper in the project does not unlock the shared framework or instrument access. A framework change would require its own authorization and scoped access.
      4. Execution and observation: after approval, it writes the helper within scope, runs only approved non-hardware checks, and records the change, approval, and actual results.
      5. Evaluation and feedback: the reviewer checks the conversion against the agreed requirement. The user performs any required hardware validation and returns findings. The agent revises within scope or stops and reports the next decision it needs.

      The guardrail checks the proposal; the person authorizes an exception; permissions constrain execution; logs record the outcome; evaluation judges whether the result meets the requirement. The harness brings these together around the model’s repeated calls.

      Relating this to the code and run below

      The runnable demonstration below uses a document lookup task, not this DUT framework. Its ContextPolicy corresponds to selecting the documentation and decisions the model sees; ToolRegistry corresponds to the actions the coding agent can call; a Hook illustrates checking a proposed action before execution; and the caps bound the loop. A veto hook alone is not a human approval workflow: that also needs a pause, a recorded decision, and a way to resume within scope. The trace illustrates observability. The DUT diagram describes how these responsibilities would apply to your project workflow, without claiming the demonstration implements its controls.

      You do not need every technique on the site to start this workflow. Reading repository files does not by itself establish a RAG system, a Markdown instruction file is not automatically a packaged skill, and using Python APIs does not require MCP. Those are separate choices if retrieval, reusable procedures, or external tool connections become necessary.

      Optional: inspect the implementation trace

      This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

      The agent harness

      The model picks each action; the harness around it decides what the model can see, what may run, and when to stop.

      Level 5 · Agent loops
      QuestionQuestionMODELmodel picks an actionmodel picks an actionTOOLsearch("DW-300 warranty")search("DW-300warranty")context policy trims old tool resultscontext policy trimsold tool resultsMODELmodel tries a second actionmodel tries asecond actionhook vetoes the callhook vetoes the callMODELmodel sees the denial and stopsmodel sees thedenial and stopsAnswerAnswerQuestionQuestionMODELmodel picks an actionmodel picks an actionTOOLsearch("DW-300 warranty")search("DW-300warranty")context policy trims old tool resultscontext policy trimsold tool resultsMODELmodel tries a second actionmodel tries asecond actionhook vetoes the callhook vetoes the callMODELmodel sees the denial and stopsmodel sees thedenial and stopsAnswerAnswer
      0of 1 step so far chosen by the model
      your code chose this stepthe model chose this step

      The run, step by step

      This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

      STEP 01 / 07Your code chose

      The question arrives

      "What does the DW-300's drain pump cost, and
      how long is it under warranty?"
      0 tokens · 0 ms

      Practical guidance

      If you use a chat app and will never run an agent, skip this page. A harness is the code wrapped around the model, written by whoever built the agent product, and there is no box for you to type in. The pages that are yours are coding agents and always-on assistants.

      If you do operate an agent product, one thing here earns your time: when an agent behaves badly, the fix is usually a setting rather than a better prompt. Four settings, and what each one looks like when it is the cause.

      What it remembers. Anthropic’s Claude Agent SDK documents automatic compaction firing as a long session grows: “When the context window approaches its limit, the SDK automatically compacts the conversation: it summarizes older history to free space, keeping your most recent exchanges and key decisions intact”[2]. Anthropic’s engineering blog describes tool result clearing, a lighter version of the same idea, as one of the “safest lightest touch forms of compaction”[3]. An agent that dropped your constraint two hours into a session did not ignore it; it summarized it away. Restate the constraint in your next message rather than starting the whole task again.

      What it may run without asking. OpenAI’s Codex documentation states the split: “Sandboxing and approvals are different controls that work together. The sandbox defines technical boundaries. The approval policy decides when the agent must stop and ask before crossing them”[6], enforced on macOS “using the built-in Seatbelt framework”[6]. Between asking every time and never asking, a policy can “keep specific approval prompt categories interactive while automatically rejecting others”[7]. Start on the setting that asks, and loosen one category at a time once you have watched what that category actually does.

      What it costs before it stops. A cap on steps or spend is a harness decision, not a model one, and a run that hits one usually ends mid-task with no error: see single agent.

      What a session hands on. A subagent may explore at length and return “only a condensed, distilled summary of its work”[3] to the harness that spawned it. That is why an agent’s account of what it did can be thinner than what it did, and why memory is a separate setting from the summary.

      Read those four in your own product’s documentation before concluding a rough session was the model’s fault.

      Implementation details

      The document lookup demonstration is the same loop single agent runs — act, check a cap, repeat: rebuilt so four moving parts are arguments to run instead of fixed in the function body: a ToolRegistry (the definitions the model is shown, and the allowed set checked before any of them runs: the split function calling makes for one call, made reusable), a ContextPolicy (a function from the growing message history to whatever the next request sends: context engineering applied inside the loop rather than once before it), a Hook (a chance to veto a call the model already chose, before the registry runs it: a silent version of what human approval does out loud), and the step and token caps.

      trim_to_budget is the context policy worth reading closely. It keeps every message except tool results, and keeps only as many of the most recent tool results as fit under a token budget, replacing older ones with a short placeholder rather than deleting them silently: a small version of what Anthropic calls tool result clearing[3]:

      examples/agent_harness/run.py · lines 75–112
      def trim_to_budget(budget_tokens: int) -> ContextPolicy:
          """The tight policy: keeps every non-tool-result message, and as many of the most recent
          tool results as fit under `budget_tokens`, dropping older ones first. Real context policies
          trim the same way -- see this page's Use it lane for how Anthropic describes tool result
          clearing and compaction -- this one trims by a plain token count to keep the point readable
          in a few lines.
      
          A tool result is recognized by its role, `tool`, never by how its text opens: the question is
          a user message, so a policy that matched on text could drop a question that happened to begin
          "Result of ...", leaving the model answering something it can no longer see. A trimmed result
          keeps its call id, because every tool call in the history still needs an answer."""
      
          def policy(messages: list[Message]) -> list[Message]:
              result_idx = [i for i, m in enumerate(messages) if _is_tool_result(m)]
              kept: set[int] = set()
              used = 0
              for i in reversed(result_idx):
                  cost = count_tokens(content_text(messages[i].content))
                  if used + cost > budget_tokens:
                      break
                  used += cost
                  kept.add(i)
              out = []
              for i, m in enumerate(messages):
                  if i in result_idx and i not in kept:
                      out.append(
                          Message(
                              role="tool",
                              content="[earlier tool result trimmed by the context policy]",
                              tool_call_id=m.tool_call_id,
                              tool_name=m.tool_name,
                          )
                      )
                  else:
                      out.append(m)
              return out
      
          return policy

      It runs fresh on every model call, not once at the start, which is what lets the two runs below diverge partway through instead of only at the first prompt. It also leaves the opening request alone: it recognizes a tool result by how the text opens, so without that guard a question starting the same way would be trimmed and the model would be answering something it could no longer see. deny_after is the hook worth reading next, three lines, and the point is that it runs after the model has already decided:

      examples/agent_harness/run.py · lines 129–142
      def deny_after(allowed_calls: int) -> Hook:
          """A hook for demonstration and testing: allows the first `allowed_calls` tool calls the
          model attempts, vetoes every one after. A real hook would read the call's own name and
          arguments; this one only counts, to keep the point -- a hook can block an action the model
          already decided to take -- in three lines."""
          seen = {"n": 0}
      
          def hook(call: ToolCall) -> tuple[bool, str]:
              seen["n"] += 1
              if seen["n"] > allowed_calls:
                  return False, f"tool budget of {allowed_calls} call(s) already spent"
              return True, ""
      
          return hook

      A real hook would look at the call’s own name and arguments instead of just counting; running one inside a sandbox that actually isolates what a tool may touch, rather than a check like this one, is code execution‘s territory, and what gets loaded into the model’s instructions in the first place (which skill, not just which tool) is skills’.

      run is the loop these parts plug into: ask the model through whatever the context policy currently allows it to see, and if it calls tools, check each one against the hook and the registry’s allowlist before running it, record what happened, and go around again until the model stops or a cap does.

      examples/agent_harness/run.py · lines 145–197
      def run(
          question: str,
          model: Model,
          embedder: Embedder | None,
          tracer: Tracer,
          *,
          corpus_dir=DEFAULT_CORPUS_DIR,
          registry: ToolRegistry = DEFAULT_REGISTRY,
          context_policy: ContextPolicy = keep_everything,
          hook: Hook = allow_everything,
          max_steps: int = MAX_STEPS,
          max_tokens: int = MAX_TOKENS,
      ) -> Answer:
          del embedder  # this harness retrieves through its tools, not a vector index
          sections = load_sections(corpus_dir)
          messages = [Message(role="system", content=SYSTEM), Message(role="user", content=question)]
          citations: list[str] = []
          tokens_used = 0
      
          for _ in range(max_steps):
              completion = model.complete(context_policy(messages), tools=registry.definitions, max_tokens=400)
              tokens_used += completion.tokens_in + completion.tokens_out
      
              if not completion.tool_calls:
                  record_completion(tracer, decided_by="model", title="Model stops and answers", completion=completion)
                  return Answer.from_text(completion.text, retrieved_sources=citations)
      
              calls_desc = ", ".join(f"{c.name}({c.arguments})" for c in completion.tool_calls)
              record_completion(tracer, decided_by="model", title="Model picks an action", completion=completion, detail=calls_desc)
              turn, calls = assistant_turn(completion, len(messages))
              messages.append(turn)
      
              for call in calls:
                  allowed, reason = hook(call)
                  if not allowed:
                      tracer.record(kind="code", decided_by="code", title="Hook vetoes the call", detail=reason)
                      messages.append(tool_result(call, f"Denied: {reason}"))
                      continue
                  if call.name not in registry.allowed:
                      result_text, cites = toolkit.unknown_tool(call.name)
                  else:
                      result_text, cites = registry.call(call, sections)
                  citations.extend(cites)
                  tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])
                  messages.append(tool_result(call, result_text))
      
              if tokens_used >= max_tokens:
                  reason = f"token budget reached: {tokens_used} >= {max_tokens}"
                  final = force_final(context_policy(messages), model, tracer, reason=reason, max_tokens=400)
                  return Answer.from_text(final.text, retrieved_sources=citations)
      
          final = force_final(context_policy(messages), model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
          return Answer.from_text(final.text, retrieved_sources=citations)

      Every tool call, its arguments, and the decision to stop are decided_by: "model"; running a tool, a hook’s veto, and forcing a final answer when a cap is reached are always decided_by: "code": the same split single_agent’s example makes, with two more kinds of code-decided step than that one has.

      tests/test_example_agent_harness.py runs the same scripted model twice with only the context policy changed, and the two runs answer differently: confidently citing the warranty term under a generous policy, saying it could not confirm the term under a tight one. Nothing about the model’s own logic changed between the two runs; only what the harness let it see did. Read that for what it is: a scripted stand-in, written to answer from whatever the harness left in front of it, so what the test proves is the mechanism, not a measurement of how much a real model’s answers move. The size of that effect is what the eval below is for, and no run of it exists yet. The same file scripts a hook that vetoes a call the model already committed to, and checks the veto shows up in the trace as the harness’s own decision, and a tool the registry advertises but will not run, and checks it fails exactly the way an unknown tool does.

      Run it yourself:

      examples/agent_harness/README.md · lines 15–15
      python -m examples.agent_harness --model stub:scripted
      When you do not need this

      Try single agent first if you have not seen the basic loop yet: this page assumes you have, and is about what surrounds it, not the loop itself.

      Move to thinking about the harness deliberately once an agent runs past a one-off demo: choosing the caps, the allowlist, the context policy and the approval settings on purpose is what turns a loop that happens to work into a system somebody can operate and debug: see safety, privacy and governance for testing one before trusting it with anything real, and operations for running one after that.

      Failure modes

      A trimmed tool result leaves a silent gap

      How to notice it
      The final answer is missing a fact an earlier tool call actually returned, with no error and no retry: the context policy dropped it before a later call, and nothing downstream says so.
      How to test for it
      Run the same scripted model through a generous context policy and a tight one on the same question and compare the final text word for word; this page's own tests do exactly this.

      A hook veto reads as the model refusing

      How to notice it
      A run stops short of an action, and it reads, from the transcript alone, like the model chose caution, when a hook actually blocked a call the model had already decided to make.
      How to test for it
      Read the trace, not the transcript. A veto is its own decided_by: "code" step; a model declining on its own is decided_by: "model". Confusing the two hides who is actually setting the policy.

      A tool the model can see is one the registry will not run

      How to notice it
      The model calls a tool whose definition it was shown, and the call fails the way an unregistered name would, because the tool was advertised but never added to the allowlist that actually runs it.
      How to test for it
      Give the model a tool definition with no matching entry in the registry's allowlist and confirm the failure looks exactly like an unknown tool, not a special error: a mismatched allowlist should never be distinguishable from a typo.

      Caps tuned for a different task cut every run short

      How to notice it
      Every run in a batch hits the step or token cap and returns a forced, partial answer, and the task looks fine in isolation: the caps were copied from a shorter task and never re-tuned.
      How to test for it
      Force a low cap on a task that genuinely needs more steps and confirm the forced answer is visibly marked, not indistinguishable from a real stop: this page's own tests do exactly this.

      Logging the model's output is not logging the harness's decisions

      How to notice it
      A run goes wrong and the only record is what the model said (not which cap fired, what a policy trimmed, or which hook denied a call) so nobody can tell whether the model or the harness caused it.
      How to test for it
      Read what actually gets logged for one run end to end and check whether a cap, a trim, or a veto shows up in it at all, or only the text the model produced.

      Cost and latency

      Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

      2Model calls, best case (act, tool, stop)
      5Model calls, worst case (step cap reached)
      ~180–310Tokens in, one action round
      ~0.6sWall time, one round trip
      Compared with single agent (level 5, the same tools)The model-call shape is identical to single agent; a harness configuration changes how much of the growing context each call actually sees, not how many calls happen. A tight context policy can cost fewer tokens per call at the same step count, at the cost of what the model can still recall.

      How to Evaluate It

      60 questionslookupmulti-hopnumericunanswerableconflicting sources

      agent_harness answers a question about the documents and cites what it used, the same task single_agent is scored on, so scripts/eval_run.py scores it against the site’s own 60-question set: run python scripts/eval_run.py --example agent_harness --model <spec> --dry to project the cost of a full run before spending anything. What is specific to this technique is running the same 60 questions through two harness configurations and comparing the two result files directly: a gap in citation hit rate or score between a generous context policy and a tight one is not model variance, it is what the harness cost the run.

      Run it

      What to monitor

      Which cap ends a run (step, token, or a real stop), how often a hook vetoes a call the model chose, and how much of the growing context a policy actually keeps versus trims on a typical run: numbers a dashboard showing only the final answer will never surface.

      Cost at volume

      Tracks the same thing single agent's does (how many actions a question actually needs) plus one more: a tighter context policy costs fewer tokens per call at any given step count, so two harnesses running the identical loop can differ in spend without differing in step count at all.

      How it fails in production

      A context policy trims a fact a later step still needed, and the run finishes anyway with a plausible but wrong answer; or a hook denies a call quietly enough that a person reading only the final text never learns anything was blocked.

      What to log

      Which cap fired if any, every hook decision and its reason, what a context policy actually dropped on each call, and the tool registry's allowlist at the time of the run: a harness that only logs the model's own output cannot be debugged when the harness itself is what went wrong.

      Try it

      1. Use it

        Pick two agent products (a coding agent and a research or deep-research tool, say) and try to name, for each, its step or turn cap, whether it asks before an irreversible action, and what it says when it hits a limit. That is the harness, not the model, and most products document at least part of it.

      2. Build it

        Run python -m examples.agent_harness --model stub:scripted from the repo root. The same model searches, looks up the part, and stops with a priced, warranty-scoped answer, all of it the harness's doing rather than the model's. Then open tests/test_example_agent_harness.py and change trim_to_budget(15) to a much larger number in the harness-changes-the-outcome test. Does the answer come back the same as the generous policy's?

      3. Either lane

        Pick one of the failure modes above and try to cause it on purpose, using the pattern in tests/test_example_agent_harness.py.

      4. Build it

        Read READ_ONLY_HEADERS and is_read_only in examples/common/bench.py, then SafetyEnvelope and Approval right after them. That is the same allowlist-plus-hook shape this page's own ToolRegistry and Hook build, in a setting where the one class of command that actually energizes a board needs a person's approval naming the set point before it runs at all.

      How it connects

      Before, after and instead of this

      Read first

      Decoded in

      Optional: products, tools, and models

      Examples of agent harnesses

      All names on this page

      11 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

      Includes ready-made harnesses, frameworks and runtimes for building your own, and products with a harness inside.

      Explore 5 more examples
      • LangGraph Tool or framework · LangChain

        A lower-level runtime for building a custom harness with persisted state, durable execution, and human checkpoints.

        Checked 09/19/2026
      • Agent Development Kit Tool or framework · Google

        A framework for building agents with tools, session state, and multi-agent coordination.

        Checked 09/19/2026
      • Microsoft Agent Framework Tool or framework · Microsoft

        Microsoft’s framework for building agents and coordinating multi-agent workflows in Python and .NET.

        Checked 09/19/2026
      • CrewAI Tool or framework · CrewAI

        A framework for coordinating agents, tasks, and flows. It supplies orchestration pieces for a custom agent system.

        Checked 09/19/2026
      • Pydantic AI Tool or framework · Pydantic

        A typed Python agent framework: define tools and validated outputs, then build the surrounding application policies.

        Checked 09/19/2026
      In practice

      Run a coding task within boundaries

      The model proposes a command. The harness checks permission, runs it in a sandbox, records the result, and builds the next request.

      Out there

      Named products, tools and models

      Products2
      • Claude CodeAnthropic · coding agent
      • CodexOpenAI · coding agent
      Tools9
      • Agent Development KitGoogle · agent framework
      • Claude Agent SDKAnthropic · agent framework
      • CrewAICrewAI · multi-agent framework
      • Deep AgentsLangChain · agent harness
      • LangGraphLangChain · graph framework
      • Microsoft Agent FrameworkMicrosoft · multi-agent framework
      • OpenAI Agents SDKOpenAI · agent framework
      • Pydantic AIPydantic · agent framework
      • Strands AgentsStrands Agents · agent harness SDK

      Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

      Where this comes from

      Primary sources

      1. Building effective agents · Anthropic, 12/19/2024 (accessed 09/19/2026)
      2. How the agent loop works · Anthropic (Claude Agent SDK documentation) (accessed 09/19/2026)
      3. Effective context engineering for AI agents · Anthropic (Engineering blog) (accessed 09/19/2026)
      4. Loop engineering: getting started with loops · Anthropic (Claude blog) (accessed 09/19/2026)
      5. Running agents · OpenAI (Agents SDK documentation) (accessed 09/19/2026)
      6. Sandboxing · OpenAI (Codex documentation) (accessed 09/19/2026)
      7. Agent approvals & security · OpenAI (Codex documentation) (accessed 09/19/2026)
      8. How Claude Code works · Anthropic (Claude Code documentation) (accessed 09/19/2026)

      Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page