# Agent graphs

_Level 06 · Teams of Agents · sourced_

Describing a team of agents and how work passes between them.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a task through a network of agent roles and explicit handoffs. Inspect which state is shared, who chooses the next route, and how the system handles a return to an earlier stage.

**Assumptions:** A role diagram is not an execution policy. State ownership and transition conditions need to be specified separately.

**Design choices:** Use agent nodes where adaptive reasoning helps and deterministic nodes for reliable checks. Add roles only when the division makes the work clearer or better.

**Request:** Investigate an incident and prepare a reviewed remediation plan.

**Starting evidence:** Roles: investigator, planner, reviewer. Mock evidence identifies inconsistency in an affected cache scope. A proposed invalidation still needs impact review. Human controls deployment.

**Action and control:** Hand evidence to planning, then a concrete proposal to review; rejection returns bounded feedback.

**Stage records (authored, not executed):**

### Input record

Roles: investigator, planner, reviewer. Mock evidence identifies inconsistency in an affected cache scope. A proposed invalidation still needs impact review. Human controls deployment.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use agent nodes where adaptive reasoning helps and deterministic nodes for reliable checks. Add roles only when the division makes the work clearer or better.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Hand evidence to planning, then a concrete proposal to review; rejection returns bounded feedback.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Proposal: invalidate the affected cache scope after approval. No production action occurs.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Role graph, handoff packet, reviewer rejection, bounded retry, and a human-approved remediation plan.

If the result falls short:
On a failed handoff, retain the originating evidence and return to the responsible node. Bound cycles so repeated review does not become endless work.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to investigations or collaborative production. Choose topology around dependencies and authority rather than modeling an organization chart for its own sake.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Proposal: invalidate the affected cache scope after approval. No production action occurs.

**Change something — Reviewer rejects an unsafe broad flush twice:** Retry cap reached: escalate with evidence and unresolved concerns. Handoffs do not expand authority.

**Decision:** Does a reviewer handoff authorize deployment?

**Answer:** No; deployment approval is separate.

**Why:** A handoff carries state and authority boundaries; prevent endless cycles and uncontrolled privilege transfer.

**Review criteria:** Role graph, handoff packet, reviewer rejection, bounded retry, and a human-approved remediation plan.

**Recovery:** On a failed handoff, retain the originating evidence and return to the responsible node. Bound cycles so repeated review does not become endless work.

**Adapt it:** Apply this to investigations or collaborative production. Choose topology around dependencies and authority rather than modeling an organization chart for its own sake.

Agent graphs is level 6 read as a graph: nodes are agents, not fixed steps; edges are handoffs
between them. State sharing and checkpointing depend on the implementation; they are not automatic properties of an agent graph. It is the third
page in the site's [graph engineering thread](/gradient_ascent/threads/graph-engineering/), after
[knowledge graphs](/gradient_ascent/techniques/knowledge-graphs/), which connect information, and
[workflow graphs](/gradient_ascent/techniques/workflow-graphs/), which connect work in code.

The useful distinction is how work is delegated and who chooses subsequent actions. A workflow can also use model-based classification or contain an agent. In the supervisor example here, one node uses a model, and its output selects
which node runs next from an explicit list of names. OpenAI's Agents SDK documents the same
arrangement: "If you have multiple possible destinations, register one handoff per destination and
let the model choose among them."[1]
Nothing else about the shape changes: code still runs every node, still checkpoints state after
each one, and a hop cap still stops the graph the model cannot see past.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 6 · Agent graphs. The same steps are described in the sections below._

## Practical guidance

There is nothing on this page for you to turn on. Agent graphs run inside a product's own
orchestration layer, assembled by whoever built it, and no product offers one as a feature with a
name and a settings page. If you are choosing a product rather than writing one, the page you want
is [always-on assistants](/gradient_ascent/techniques/agent-teammates/), and
[a lead and its workers](/gradient_ascent/techniques/orchestrator-workers/) is the version of
this shape you can actually watch happen.

The names here belong to developers. LangGraph, Microsoft Agent Framework, CrewAI and Google's
Agent Development Kit are all registered against this technique, and an open protocol exists for
the case where agents built on different ones of them have to hand work to each other. Agent2Agent
(A2A) states its purpose as "Connect agents built on different platforms (LangGraph, CrewAI,
Semantic Kernel, custom solutions) to create powerful, composite AI systems."[2] Its own
site says A2A was "Originally developed by Google and now donated to the Linux
Foundation"[2], and is at version 1.0. Microsoft's own framework documentation names the
shape plainly: "graph-based workflows supporting sequential, concurrent, handoff, and group
collaboration patterns; includes checkpointing, streaming, human-in-the-loop, and
time-travel"[3]: checkpointing and handoff, named together, are exactly the two things
that separate a graph of agents from a plain loop.

One consequence does reach you, and it is the only thing here to act on. When a multi-step
assistant's tone or accuracy changes partway through one answer, a handoff happened. Ask the
vendor whether the product labels them: "When more than one agent works on a request, does the
answer or the log say which one produced which part?" A product that tells you gives you something
to check. A product where the handoff is invisible gives you an answer whose author you cannot
identify, which is worth weighing before you buy, not after a wrong answer you cannot trace.

## Implementation details

The example extends the idea in [workflow graphs](/gradient_ascent/techniques/workflow-graphs/)'
own runner without editing that file: nodes are plain functions over shared state, and the runner
checkpoints after every node. The only change is that one node, the supervisor, is not a fixed
code rule; it is a model call, and the loop acts on whatever node name that call's output names.

`examples/agent_graphs/run.py` (lines 82-120)

```python
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    max_research_hops: int = MAX_RESEARCH_HOPS,
) -> Answer:
    del embedder  # retrieval here is keyword search, like workflow_graphs
    sections = load_sections(corpus_dir)
    findings: list[tuple[str, str]] = []
    hops = 0

    while True:
        if hops >= max_research_hops:
            tracer.record(
                kind="code", decided_by="code", title="Hop cap reached",
                detail=f"{hops} research hops >= {max_research_hops}; forcing write",
            )
            choice = "write"
        else:
            choice = _supervisor_choose(question, findings, model, tracer)
            if choice not in ALLOWED_HANDOFFS:
                tracer.record(
                    kind="code", decided_by="code", title="Handoff blocked",
                    detail=f"{choice!r} is not in the allowlist {sorted(ALLOWED_HANDOFFS)}; forcing write",
                )
                choice = "write"

        if choice == "write":
            answer_text = _node_write(question, findings, model, tracer)
            break

        hops += 1
        _node_research(question, sections, findings, tracer)

    citations = sorted({cite for cite, _ in findings})
    return Answer.from_text(answer_text, retrieved_sources=citations)
```

Code never hands that name straight to a node without checking it first. `ALLOWED_HANDOFFS` plays
the role OpenAI's own `is_enabled` handoff switch does: "a boolean or a function that returns a
boolean, allowing you to dynamically enable or disable the handoff at runtime"[1], here
fixed rather than dynamic: a name the model returns that is not `research` or `write` is caught
before the graph acts on it, recorded as a `Handoff blocked` step, and the run is forced to `write`
instead. `tests/test_example_agent_graphs.py` scripts exactly this: the model answers
`delete_database`, and the test checks the run never treats that as a research hop, cites nothing,
and still returns an answer rather than failing.

`research` and `write` are ordinary code, the same as any node in a workflow graph. `research`
re-searches the whole corpus for the original question, skipping any section already found, so a
second hop turns up the next-best match instead of repeating the first: the reason this run needs
three hops to gather the manual's figure, the bulletin that revises it, and the revised number
itself, in that order. `MAX_RESEARCH_HOPS` (4 by default) caps how many times the supervisor may
send the team back to `research`; past the cap, code forces `write` on its own without asking the
model again, unlike a blocked handoff, which does still get one more supervisor call afterward.

Run it yourself:

`examples/agent_graphs/README.md` (lines 17-17)

```text
python -m examples.agent_graphs --model stub:scripted
```

Every step the trace names `Supervisor picks the next agent` is `decided_by: "model"`; every checkpoint, the
blocked-handoff step and the hop-cap step are `decided_by: "code"`. This is the same distinction [workflow graphs](/gradient_ascent/techniques/workflow-graphs/)' own page draws, with one more place a
model, not code, now gets to choose.

## When you do not need this

Try [workflow graphs](/gradient_ascent/techniques/workflow-graphs/) first if every transition
rule can be written down before the graph runs: most graphs can, and a fixed rule costs nothing
to evaluate and is wrong the same way every time it is wrong.

Try a [single agent](/gradient_ascent/techniques/single-agent/) instead if one model, one loop and
one context window is enough: a graph of separate agents is only worth its extra machinery once
the task needs more than one agent's own context to hold.

Move up to agent graphs once a node's own output has to pick which agent runs next, not just what
a fixed step does. This is the same reason [workflow graphs](/gradient_ascent/techniques/workflow-graphs/)
justifies moving up from a plain chain, one level higher.

## Failure modes

### A handoff outside the allowlist

- **How to notice it:** The supervisor's output names something that is not a real node (a slightly different word, or something invented outright), and unless it is caught, the graph either fails trying to run a node that does not exist or silently falls through to whatever the code happens to do by default.
- **How to test for it:** Script the supervisor to return a name outside the allowlist (this page's own test suite does exactly this) and confirm the run is forced to a safe node instead of failing or quietly continuing as if nothing happened.

### The supervisor never converges

- **How to notice it:** The supervisor keeps sending the team back to research, finding less and less that is new each time, until the hop cap forces a stop rather than the supervisor choosing to stop on its own.
- **How to test for it:** Read what each hop's checkpoint actually added to the findings. Real progress narrows toward an answer; a stalled supervisor keeps asking for the same kind of information a later hop already supplied.

### A node reads state a checkpoint never wrote

- **How to notice it:** A node expects a field in the shared state that no earlier node actually set, so it either fails or silently treats it as empty, and the next handoff is decided on less information than the run actually gathered.
- **How to test for it:** Compare every checkpoint's own record of what it wrote against what the next node reads. A field read but never written by anything upstream is this failure.

### The hop cap ships a thin answer

- **How to notice it:** The cap is reached before the supervisor chose to write on its own, and the write node drafts from whatever partial findings exist, silently unless the 'Hop cap reached' step is surfaced somewhere a person can see it.
- **How to test for it:** Script a supervisor that always answers 'research' (this page's own test suite does exactly this) and confirm the run still returns an answer once the cap is hit, and that the trace says the cap forced it.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, best case (supervisor writes immediately):** 2
- **Model calls, this run (3 research hops, then write):** 5
- **Tokens in, one supervisor call:** ~140–210
- **Wall time, one hop (handoff plus checkpoint):** ~0.3s

**Compared with workflow graphs, the same two nodes with every edge fixed in code (level 3).** A workflow graph pays for exactly the nodes its code visits, every run, whether or not that is enough. This run cost five calls because the question needed three research hops; a simpler question through the same graph costs two, and a workflow graph could not tell the difference in advance.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

`agent_graphs` answers the same question-about-the-documents task `rag` and `single_agent` are
scored on (the write node's final answer cites the sections the research node actually found), so
it fits the site's 60-question set the same way: exact or rubric match, citation hit rate, plus a
number specific to this shape: hops used per question against the cap, which shows directly
whether a question needed one pass through the graph or several.

`scripts/eval_run.py` counts `agent_graphs` among the examples the question set can score (see
`docs/EVALS.md`). No result file exists for it yet, so this page cannot say a number for any of
it. Run `python scripts/eval_run.py --example agent_graphs --model <spec> --dry` to project the
cost of a real run first; that projection assumes the full hop cap, which is the ceiling, not
what a typical question costs.

## Run it

**What to monitor.** Hops used per question against the cap, the share of runs where a handoff was blocked as outside the allowlist, and how often the hop cap fires before the supervisor chooses to write on its own.

**Cost at volume.** Cost tracks hops, not a fixed step count: a question the graph resolves in one research hop costs two calls, and one that needs the full hop cap costs several times that, for the same question, the way a single agent's cost tracks how many actions it takes.

**How it fails in production.** The supervisor keeps handing off to research without making progress, silently spending the whole hop cap on a question the graph's two agents were never going to resolve, or a handoff target the model invented gets blocked and the run ships a thinner answer than the question needed.

**What to log.** Every supervisor decision with its full output text, every checkpoint's state, any blocked-handoff event, and which cap (if any) forced the stop, so a bad answer traces back to a specific handoff rather than an unexplained partial result.

## Try it

1. **Use it.** Watch a multi-step assistant handle a task that plausibly needs more than one kind of expertise (research plus writing, say). Does anything in its reply say a different specialist or stage handled part of it, or is the switch invisible?
2. **Build it.** Run python -m examples.agent_graphs --model stub:scripted from the repo root. The supervisor sends the team to research three times, each hop surfacing a section the last one did not, then hands off to write, and the answer cites the service bulletin that supersedes the manual. Run it again with --model stub to see why the scripted sequence exists: the echo is not one of the two agent names, so the first hop goes nowhere and the run ends with no citations at all.
3. **Either lane.** Pick one of the failure modes above and try to script a StubModel response that causes it on purpose, using the pattern in tests/test_example_agent_graphs.py.


## Sources

1. [Handoffs](https://openai.github.io/openai-agents-python/handoffs/) — OpenAI (Agents SDK documentation) (accessed 2026-09-19)
2. [Agent2Agent (A2A) Protocol](https://a2a-protocol.org/v1.0.0/) — Linux Foundation (Agent2Agent Protocol) (accessed 2026-09-19)
3. [microsoft/agent-framework](https://github.com/microsoft/agent-framework) — Microsoft (accessed 2026-09-19)


Last reviewed 2026-09-19.
