# Workflow graphs

_Level 03 · Workflows · sourced_

Describing a workflow as steps and the connections between them.


## Try this in a recipe
- [Build a weekly update without invented progress](/gradient_ascent/recipes/weekly-status-report.md): Extract evidence into a checked table, then draft an update from that table in a fixed two-call workflow.

## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a task through explicit branches, handoffs, and a review point. The example makes the route visible so you can inspect what happens when information is missing or a step must be repeated.

**Assumptions:** Transitions need defined conditions and state. A diagram alone does not ensure that a resumed or retried run behaves correctly.

**Design choices:** Use deterministic transitions for known business rules, and model calls for work needing language judgment. Include human review where the actual consequences justify it.

**Request:** Prepare this week's portfolio report and send only after I approve its content and recipients.

**Starting evidence:** Fictional week W12: Atlas delayed to Friday, Beacon complete, Cedar has no fresh evidence. Recipients: project leads.

**Action and control:** Collect → reconcile → draft → review → approved-version delivery. Missing evidence stays unresolved.

**Stage records (authored, not executed):**

### W12 evidence register

Atlas tracker: delivery delayed to Friday.
Beacon notes: work complete.
Cedar: no fresh evidence.
Previous report: continuity only, not proof of current status.
Proposed audience: project leads.

What changed: The current period starts with two supported updates and one explicit gap.

### State and transitions

Collect → reconcile → draft → review → delivery.
Missing update → retain an unknown status.
Edited report or recipients → return to review.
Uncertain delivery → reconcile before retrying.
Policy in this example: human approves distribution.

What changed: Branches distinguish missing evidence, changed approval scope, and uncertain external effects.

### Reconciled working record

Atlas: schedule risk; source = tracker.
Beacon: complete; source = notes.
Cedar: unknown; source = missing.
Unsupported inference removed: Cedar is green because no problem was reported.
Draft state: ready for review, not approved.

What changed: Collection has become a reportable status record without inventing Cedar's progress.

### W12 report · v1

Atlas schedule risk; Beacon complete; Cedar unknown.
Evidence: Atlas tracker; Beacon notes; Cedar missing.
Recipients: project leads.
Distribution: pending approval of this content and audience.

What changed: The controls below model approval and delivery separately. Editing either field invalidates approval.

### Recovery record

Illustrated alternate event: send times out after possible acceptance.
Known: a send was attempted.
Unknown: whether the destination accepted it.
Recovery: inspect delivery status using the same report identity before resending.
Production duplicate protection is not implemented by this page.

What changed: The timeout branch is a written teaching record; the local send button only demonstrates an in-memory gate.

### Adaptation plan

Replace: projects, reporting period, connectors, and audience.
Choose: which missing updates block a draft and which can be marked unknown.
Choose: review rules for distribution in your organization.
Keep: evidence freshness, explicit state, and safe handling of uncertain delivery.

What changed: A personal draft may need no approval gate; a shared report may need one at distribution.

**Sample result:** Draft v1: Atlas schedule risk; Beacon complete; Cedar unknown. Sources: Atlas tracker, Beacon notes, Cedar missing. Audience: project leads.

**Change something — A send times out after possible acceptance:** Delivery uncertain. Reconcile the sent record and reuse the report identifier; do not blindly issue another send.

**Decision:** Should a timed-out send be retried as a new report?

**Answer:** No; reconcile and protect against duplicates.

**Why:** Missing data, connector failures, edits after approval, and send retries need separate states; a scheduled fixed workflow is not automatically an autonomous agent.

**Review criteria:** A collection-to-review state graph, evidence links, unresolved-items queue, version-bound approval, and simulated duplicate-safe delivery.

**Recovery:** Resume from a recorded state, invalidate downstream results when their inputs change, and distinguish safe retries from duplicate external actions.

**Adapt it:** Replace report generation with another repeatable process. Choose your own sources, branches, review policy, and completion criteria; not every workflow needs every gate shown here.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a task through explicit branches, handoffs, and a review point. The example makes the route visible so you can inspect what happens when information is missing or a step must be repeated.

**Assumptions:** Transitions need defined conditions and state. A diagram alone does not ensure that a resumed or retried run behaves correctly.

**Design choices:** Use deterministic transitions for known business rules, and model calls for work needing language judgment. Include human review where the actual consequences justify it.

**Request:** Organize shared household purchases with review before checkout.

**Starting evidence:** List: rice, soap. One item unavailable. Spending cap: $40. No substitute preapproved.

**Action and control:** Collect requests → check availability → propose alternatives → review basket → simulated checkout.

**Stage records (authored, not executed):**

### Input record

List: rice, soap. One item unavailable. Spending cap: $40. No substitute preapproved.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use deterministic transitions for known business rules, and model calls for work needing language judgment. Include human review where the actual consequences justify it.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Collect requests → check availability → propose alternatives → review basket → simulated checkout.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Basket draft retains available items and asks about the unavailable item. No purchase occurs.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Inspect the exact approved basket and ensure checkout cannot precede that decision.

If the result falls short:
Resume from a recorded state, invalidate downstream results when their inputs change, and distinguish safe retries from duplicate external actions.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Replace report generation with another repeatable process. Choose your own sources, branches, review policy, and completion criteria; not every workflow needs every gate shown here.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Basket draft retains available items and asks about the unavailable item. No purchase occurs.

**Change something — Availability changes after approval:** Return the changed basket to review. The old decision does not cover a new substitute or price.

**Decision:** Should a changed basket proceed under the old approval?

**Answer:** No; review the changed contents and total.

**Why:** A workflow state tracks prerequisites and exceptions, not just a happy-path sequence.

**Review criteria:** Inspect the exact approved basket and ensure checkout cannot precede that decision.

**Recovery:** Resume from a recorded state, invalidate downstream results when their inputs change, and distinguish safe retries from duplicate external actions.

**Adapt it:** Replace report generation with another repeatable process. Choose your own sources, branches, review policy, and completion criteria; not every workflow needs every gate shown here.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a task through explicit branches, handoffs, and a review point. The example makes the route visible so you can inspect what happens when information is missing or a step must be repeated.

**Assumptions:** Transitions need defined conditions and state. A diagram alone does not ensure that a resumed or retried run behaves correctly.

**Design choices:** Use deterministic transitions for known business rules, and model calls for work needing language judgment. Include human review where the actual consequences justify it.

**Request:** Process an engineering change through impact analysis, review, and release preparation.

**Starting evidence:** Change request affects firmware and test documentation. Review requires both owners' sign-off.

**Action and control:** Create explicit states for impact assessment, parallel reviews, reconciliation, and approved release preparation.

**Stage records (authored, not executed):**

### Input record

Change request affects firmware and test documentation. Review requires both owners' sign-off.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use deterministic transitions for known business rules, and model calls for work needing language judgment. Include human review where the actual consequences justify it.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Create explicit states for impact assessment, parallel reviews, reconciliation, and approved release preparation.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Release package remains pending until both required reviews are complete. No deployment occurs.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Inspect review state, changed scope, and which approvals remain valid.

If the result falls short:
Resume from a recorded state, invalidate downstream results when their inputs change, and distinguish safe retries from duplicate external actions.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Replace report generation with another repeatable process. Choose your own sources, branches, review policy, and completion criteria; not every workflow needs every gate shown here.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Release package remains pending until both required reviews are complete. No deployment occurs.

**Change something — Documentation reviewer rejects the change:** Return affected work for revision without discarding the completed unrelated review; reassess approvals if scope changes.

**Decision:** Does one reviewer approving complete a two-owner gate?

**Answer:** No; satisfy every required review.

**Why:** Graph transitions should encode the actual review policy and revision behavior.

**Review criteria:** Inspect review state, changed scope, and which approvals remain valid.

**Recovery:** Resume from a recorded state, invalidate downstream results when their inputs change, and distinguish safe retries from duplicate external actions.

**Adapt it:** Replace report generation with another repeatable process. Choose your own sources, branches, review policy, and completion criteria; not every workflow needs every gate shown here.

A workflow graph describes a process as nodes and edges instead of nested code: each node does
one unit of work, and each edge decides what runs next. LangGraph, a framework built around the
idea, defines all three plainly. Of nodes: "Functions that encode the logic of your agents. They
receive the current state as input, perform some computation or side-effect, and return an
updated state." Of edges: "Functions that determine which Node to execute next based on the
current state." Of state: "A shared data structure that represents the current snapshot of your
application"[1]. State is passed from node to node rather than living in local
variables.

The graph describes the available steps and transitions. A transition can use a deterministic
condition or a model-produced classification; either way, the surrounding application defines
how the result selects the next step. Persistence and checkpointing are additional capabilities,
not requirements for something to be a workflow graph.

The distinction from [agent graphs](/gradient_ascent/techniques/agent-graphs/) is the work inside
the nodes and who chooses subsequent actions, not simply the presence of edges. A node can run
a fixed operation or an agent loop, and a system can combine both. The site's levels organize
these patterns for explanation; they are not universal framework categories. It is also the middle page of the site's
[graph engineering thread](/gradient_ascent/threads/graph-engineering/):
[knowledge graphs](/gradient_ascent/techniques/knowledge-graphs/) connect information, agent
graphs connect work, and a workflow graph is what you draw when the work is settled in advance.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 3 · Workflow graphs. The same steps are described in the sections below._

## Practical guidance

Build a small workflow graph in a visual automation canvas: a trigger, a sequence of steps, and
at least one branch node that sends the run one way or another depending on something in the
data so far, such as sending it one place if an amount is over a threshold and somewhere else if
it is not. Dify describes the appeal as a visible plan rather than a hidden prompt: "visual
building blocks and AI prompts to define how an app thinks, retrieves data, makes decisions, uses
tools, asks for human input, and completes tasks," turning "Prompt Logic Into a Visible Execution
Path"[2].

A branch is worth trusting only once you know what actually decides it. A rule, such as one that
checks whether a category field equals a fixed value, is fully predictable: build one input for
each path on purpose and confirm each one lands where you expect. A branch that asks a model
which way to go, inside that same node, wears the identical visual shape but behaves differently:
you cannot enumerate every path in advance the way you can with a rule, so test it with several
inputs worded differently but aimed at the same branch, not just one.

After a run finishes, open its history if the tool keeps one and read what each node received and
returned, in order, not only the final output. If a run failed or took the wrong branch, that
per-node history is where the wrong turn actually shows up. And if a run fails partway through,
check whether the tool can resume from where it stopped instead of starting the whole thing over;
that depends on whether it actually saved the state at each step along the way, which is usually
a setting worth turning on rather than something you get for free.

A graph is more machinery than a task needs when nothing in it ever branches. A plain ordered
list of steps does that job with one less thing to build, test and reread later, and is worth
reaching for first.

## Implementation details

The example is the same draft/check/revise loop as [write and check](/gradient_ascent/techniques/evaluator-optimizer/), run through a small graph
executor instead of a hand-written loop, plus one branch a single loop does not express as
cleanly: if retrieval finds nothing, the graph goes straight to a dead end (`no_match`) instead
of drafting from zero sources.

A node is a plain function of the shared state that returns an updated state, and nothing else.
This one is the whole of `retrieve`:

`examples/workflow_graphs/run.py` (lines 54-57)

```python
def _node_retrieve(state: State, sections: dict[str, Section], model: Model) -> State:
    hits = [s for s, score in bm25_search(sections, state["question"], k=RETRIEVE_K) if score > 0]
    state["source_cites"] = [s.cite for s in hits]
    return state
```

`NODES` collects five of those under their ids; `EDGES` holds one function per node, each
reading the state and returning the id of the next node, or `None` to stop. The runner is one
`while` loop: run the current node, record what it did, ask its edge function what comes next,
write a checkpoint, move on. The draft, check and revise nodes are the same
prompts [write and check](/gradient_ascent/techniques/evaluator-optimizer/) uses, so they are
left out below; the graph machinery is the part worth reading here.

`examples/workflow_graphs/run.py` (lines 95-154)

```python
NODES: dict[str, Callable[[State, dict, Model], State]] = {
    "retrieve": _node_retrieve,
    "no_match": _node_no_match,
    "draft": _node_draft,
    "check": _node_check,
    "revise": _node_revise,
}

def _edge_from_retrieve(state: State) -> str | None:
    return "draft" if state["source_cites"] else "no_match"

def _edge_from_check(state: State) -> str | None:
    if state["verdict"] == "ok" or state["revisions"] >= MAX_REVISIONS:
        return None
    return "revise"

EDGES: dict[str, Callable[[State], str | None]] = {
    "retrieve": _edge_from_retrieve,
    "no_match": lambda state: None,
    "draft": lambda state: "check",
    "check": _edge_from_check,
    "revise": lambda state: "check",
}

def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer:
    del embedder  # retrieval here is keyword search, not a vector index
    sections = load_sections(corpus_dir)
    state: State = {"question": question, "source_cites": [], "draft_text": "", "verdict": None, "revisions": 0}

    node_id: str | None = "retrieve"
    while node_id is not None:
        state = NODES[node_id](state, sections, model)
        completion = state.pop("_completion", None)
        if completion is not None:
            tracer.record(
                kind="model", decided_by="code", title=f"Node: {node_id}", detail=completion.text[:200],
                tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
            )
        else:
            tracer.record(kind="code", decided_by="code", title=f"Node: {node_id}", detail=", ".join(state["source_cites"]) or "none")
        next_id = EDGES[node_id](state)
        tracer.record(
            kind="code", decided_by="code", title=f"Checkpoint after '{node_id}'",
            detail=f"revisions={state['revisions']} verdict={state.get('verdict')!r} -> next: {next_id or 'stop'}",
        )
        node_id = next_id

    citations = cited_sources(state["draft_text"])
    return Answer(text=state["draft_text"], citations=citations, retrieved_sources=state["source_cites"])
```

The checkpoint is the concrete payoff graphs are usually sold on. `state`'s durable fields are
all plain values (strings, an int, a list of citation strings) so recording it after every node
is a real `json.dumps`, not an idea that would need a redesign to actually implement. A crashed
run could resume from the last written checkpoint by loading that same dict and re-entering the
loop at the node the checkpoint names, with no change to `NODES` or `EDGES` at all.

Whether the graph is worth it here is a fair question, and the honest answer is: barely, for
this exact example. Five nodes and a handful of edge functions do the same job
write-and-check's plain `while` loop does in fewer lines, with one extra indirection (looking a
function up in a dict instead of calling it directly) that buys nothing when there is only one
reasonable order to run things in. The place a graph earns that cost back is where this example
starts to gesture at it but does not fully need it: a genuine branch (`no_match`) that a nested
`if` inside a longer function would have buried, a state shape simple enough to checkpoint for
real, and node functions that would still make sense wired into a different graph: none of
which a straight-line function forbids, but a graph makes structurally obvious rather than
something a reader has to reconstruct from control flow.

Run it yourself:

`examples/workflow_graphs/README.md` (lines 17-17)

```text
python -m examples.workflow_graphs --model stub:scripted
```

## When you do not need this

Try a plain function with an ordinary `if`/`while`, the way [write and check](/gradient_ascent/techniques/evaluator-optimizer/)'s example does, first if the
process has one obvious order and no real branch: a graph's nodes-and-edges indirection is
ceremony when there is only one path through the code anyway.

Move up to a workflow graph once the process has a real branch that a nested `if` would bury, a
state shape worth checkpointing between steps, or node functions you expect to reuse in more
than one wiring.

## Failure modes

### An edge function with a bug routes to the wrong node silently

- **How to notice it:** The run completes and returns an answer, but a later step is missing something an earlier node actually produced, because an edge function read the wrong state field or compared it wrong.
- **How to test for it:** Unit test every edge function directly, with a small hand-built state dict for each branch it can take, the same way you would test any other pure function: no model or graph run required.

### The graph has an unreachable node

- **How to notice it:** A node exists in NODES but no edge function ever returns its id, so it is dead code that looks, from the diagram, like part of the live process.
- **How to test for it:** List every node id and confirm each one appears as at least one edge function’s return value somewhere in EDGES; one that never does is either genuinely dead or the sign of a typo in an edge function.

### A cycle with no exit condition

- **How to notice it:** Two edge functions route back and forth between the same two nodes forever, because neither one's condition can ever become the one that stops the loop.
- **How to test for it:** Trace every cycle in the edge graph by hand and confirm at least one edge function's condition is guaranteed to change monotonically (a counter that only increases, capped in code) rather than depending only on model output that might never satisfy it.

### The checkpoint is incomplete

- **How to notice it:** A resumed run behaves differently from an uninterrupted one, because some field the process actually depends on was never part of the checkpointed state (it lived in a local variable, or a node's closure) and so was lost across the resume.
- **How to test for it:** Serialize the state after each node with the real checkpoint mechanism, load it back into a fresh process, and resume from there; compare the final answer against an uninterrupted run on the same question.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Nodes visited, no revision needed:** 3
- **Model calls, no revision needed:** 2
- **Trace steps per node visited:** 2
- **Wall time, no revision needed:** ~1.4s

**Compared with write and check (level 3, same task).** Same model calls as the hand-written loop for the same outcome; the graph runner adds a checkpoint step after each node, which is bookkeeping cost in trace size and code, not in tokens spent on the model.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

The same 60 questions as every other technique, with a narrow interest in the result. Because
the graph runs the same logic as write and check, its score on citation hit rate and revision
count should match
that page's result file once one exists: a divergence between the two would point at a bug in
the graph wiring, not a difference in what the technique can do. The `no_match` branch is what
`unanswerable` questions specifically exercise: whether retrieval finding nothing correctly ends
the run at a dead end rather than drafting from an empty source list.

No result file exists yet (see `docs/EVALS.md`). Run `python scripts/eval_run.py --example
workflow_graphs --model <spec> --dry` to project the cost of a real run first.

## Run it

**What to monitor.** Which node a run stopped on and why (an edge function's return value), not just the final answer. A node that is visited far more or less often than expected is a sign a branch condition drifted from what real traffic actually looks like.

**Cost at volume.** The same as whatever the underlying nodes cost, since the graph runner itself makes no model calls of its own. The checkpoint write after every node adds a small, fixed storage cost per node visited, independent of question difficulty.

**How it fails in production.** An edge function's condition stops matching reality (a state field's shape changed upstream and the comparison silently always takes the same branch), which nothing catches unless every edge function is tested against the specific states it is meant to branch on.

**What to log.** The full state dict at every checkpoint, and which edge each transition took, not just the final node's output. A wrong final answer is only debuggable if you can see the exact state the run was in when it decided each turn.

## Try it

1. **Use it.** Find an automation tool with a visual if/else node. Read one branch condition: a fixed rule, or a model deciding?
2. **Build it.** Run python -m examples.workflow_graphs --model stub:scripted from the repo root: retrieve, draft, a check that fails on an unretrieved citation, revise, check, stop, checkpointing after every node. Now ask a question of gibberish words: retrieval returns nothing, the graph goes straight to no_match, and two checkpoints stand in for five.
3. **Either lane.** Draw a process you run as nodes and edges, then count the edges whose condition you could not write as a line of code. That is where it stops being this level.


## Sources

1. [Graph API overview](https://docs.langchain.com/oss/python/langgraph/graph-api) — LangChain (LangGraph documentation) (accessed 2026-09-19)
2. [Dify](https://dify.ai/) — LangGenius (accessed 2026-09-19)


Last reviewed 2026-09-19.
