Level 03

Workflows

Connect model calls through predefined steps, branches, checks, and retries. A model can classify an input or evaluate a result to route the workflow; software still defines the available paths.

Level 03

Who decides the next stepSoftware defines the steps and allowed branches; model outputs can select among them.
Techniques

What is at this level

Software organizes the steps

Prompt chaining

Sourced

Splitting a task into steps, each with its own prompt.

Routing

Sourced

Sorting inputs and sending each one to the right prompt.

Parallel calls

Sourced

Running several prompts at once and combining the results.

Write and check

Sourced

One prompt writes, another checks, and the loop repeats until the check passes.

Workflow graphs

Sourced

Describing a workflow as steps and the connections between them.

Human approval

Sourced

Pausing for a person to approve or correct.

Upgrade conditions

When something here is not enough

Each line names the failure that justifies moving to a higher level.

Prompt chaining → Single agent

The steps cannot be written down in advance.

Routing → Function calling

The right destination depends on something the message does not say, so the choice needs a lookup before it can be made.

Parallel calls → Lead agent and workers

The subtasks cannot be known until the task is read.

Write and check → Review and debate

One reviewer shares the author's blind spots.

Workflow graphs → Agent graphs

A node needs its own loop and tools, not one call.

Recipes

Jobs that top out here

Each one needs this level and no higher, and says why.

Level 1 + Level 3

Sort an inbox

Sorts mail into fixed categories and produces structured output. A person approves anything that gets sent. The categories are known in advance, so an agent is not needed.

This example uses level 3
↗
Level 1 + Level 3

Turn photos and PDFs into records

Reads the image or PDF, fills a fixed schema, and saves the record once a person confirms it.

This example uses level 3
↗
Level 1 + Level 3

Voice notes into structured entries

Transcribes a voice note, splits it into steps, and turns each step into a structured entry that an eval set checks for accuracy.

This example uses level 3
↗
Level 1 + Level 3

Nightly source monitor

Runs on a timer, diffs a set of public pages in code, and asks a model one question about each change. Level 3: the schedule and the checkpoint are infrastructure, not agency.

This example uses level 3
↗
Level 3

Drafting with a reviewer

One prompt drafts a piece of writing and another checks it against a rubric, repeating until the draft passes.

This example uses level 3
↗
Level 0 + Level 1 + Level 3

Match invoices to purchase orders

Extract invoice fields, then use code to match purchase orders and compare amounts. Differences go to a person; the model never decides whether the totals reconcile.

This example uses level 3
↗
Level 1 + Level 3

Check an agreement against your own checklist

Check an agreement against a fixed checklist, with cited clauses for each finding. Merge the findings for a person to review.

This example uses level 3
↗
Level 1 + Level 3

Turn an incident write-up into a runbook

Turn an incident write-up into a timeline and repeatable steps. Check owners and success criteria, then ask the incident lead to approve it.

This example uses level 3
↗
Level 1 + Level 3

Turn a script into a shot list

Split a script into scenes and shots, then check that every line is covered and every shot has a source. A person reviews the plan; drawing frames is a separate task.

This example uses level 3
↗
Level 0 + Level 1 + Level 3

Keep a tracker document current from several sources

Keep a shared tracker current through source comparisons and a review queue. Model proposals and changes to human-written fields need approval; missing evidence is flagged.

This example uses level 3
↗
Level 0 + Level 1 + Level 3

Pull an instrument's accuracy table out of its manual

Extract specification rows from a manual, validate their structure, and calculate uncertainty in code. A person verifies ranges, intervals, and conditions against the source.

This example uses level 3
↗
Level 1 + Level 3

Sort failing units and operator notes into causes

Failing measurements and free-text operator notes are sorted into the causes the failure analysis guide already lists, then routed. A person confirms before anything is scrapped or reworked. The categories are known in advance, so this is classification into fixed classes and not an agent.

This example uses level 3
↗
Level 1 + Level 3

Check a board against the design rules document

A bill of materials and a netlist summary are checked rule by rule against the written design rules. One pass drafts findings, a second checks each finding against the rule text it cites and drops the ones that cite nothing. Level 3, because code decides every step and the rules do not change between boards. This is the rule check that happens before a review meeting, not the design review report itself: for the report, and the characterization data behind it, see the two recipes this page links in its first paragraph.

This example uses level 3
↗
Level 1 + Level 3

Turn a requirements list into a test plan

A fixed chain: read the requirements, propose a test for each, build the traceability table, then check that every requirement has a test and every test names a requirement. A person approves before any of it is adopted. The order of the steps is known in advance, which is what keeps this at level 3.

This example uses level 3
↗
Out there

Named products, tools and models

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Products that work this way13
  • Claude Code subagentsAnthropic · multi-agent feature of a coding agent
  • ClineCline · open-source coding agent
  • DifyLangGenius · visual workflow builder
  • FlowiseFlowise · visual workflow builderRetired 2026-08-10
  • Grok HeavySpaceXAI · several agents answering one question
  • GumloopGumloop · visual workflow builder
  • JulesGoogle · coding agent
  • LangflowIBM · visual workflow builder
  • MakeCelonis · automation service
  • n8nn8n · automation service, self-hostable
  • Power AutomateMicrosoft · automation service
  • Replit AgentReplit · coding agent
  • ZapierZapier · automation service
Tools for building it13
  • Apache Airflowopen source · workflow engine
  • Batch APIOpenAI · asynchronous batch processing api
  • Batch APIGoogle · asynchronous batch processing api · formerly Batch Mode
  • Deep AgentsLangChain · agent harness
  • DSPyStanford NLP · prompt programs and optimizers
  • Haystackdeepset · retrieval framework
  • InngestInngest · durable workflow engine
  • LangChainLangChain · application framework
  • LangGraphLangChain · graph framework
  • Message Batches APIAnthropic · asynchronous batch processing api
  • PrefectPrefect · workflow engine
  • Semantic RouterAurelio Labs · library that routes a request by meaning
  • TemporalTemporal · durable workflow engine
Models1
  • JevTypeSafe AI · system one decision model
Frontier

What is still unsolved here

Open problems at this level, what people are trying, and the source each rests on. Read 09/19/2026. This block ages faster than the rest of the page, and nothing in it predicts which approach wins.

A wrong number introduced early in a fixed chain does not simply persist. It becomes a wrong calculation, then a sentence of prose, then a conclusion, and it gets harder to spot at every handoff, because each stage makes it look more like ordinary output.

What people are trying

Putting a check between stages rather than only at the end. One measurement through a four-stage pipeline found an early gate catches much more than the same check run once on the final result.

Where it bites: Prompt chaining

A router sorts a request into a branch before anything else runs, and routers still send plenty of requests somewhere other than where they would have been answered best. Under one benchmark that compared many of them on the same footing, several recent methods, commercial ones included, did not clearly beat a simple baseline.

What people are trying

Shared testbeds spanning many models and many routing methods, so routers are compared on the same task set instead of each paper's own. Some work scores cost and latency alongside accuracy rather than accuracy alone.

Where it bites: Routing

The check half of a write-and-check loop is usually another model call, and a model grading text written in its own style can favor it for reasons that have nothing to do with quality. Picking a more capable judge does not reliably fix it.

What people are trying

Measuring the bias directly, by grading pairs of answers built to differ in style and not in quality, so any preference shown is the bias. Structured multi-part rubrics are being tested as a way to cut it down.

Where it bites: Write and check

Putting a person in the loop is the easy part. How much they should be asked to decide, and how to keep them from waving through whatever the system already proposed, is not settled, and NIST says so in its own guidance rather than prescribing a split.

What people are trying

NIST's playbook asks teams to design the human role deliberately, account for the biases that make a reviewer agree too readily, and measure whether oversight is happening at all. Override rates and time spent per review are what tell you, rather than the presence of a checkpoint.

  • Map (AI RMF Playbook, AIRC) · NIST, Trustworthy and Responsible AI Resource Center · read 09/19/2026
    Questions remain about how to configure humans and automation for managing AI risks.

Where it bites: Human approval

← Level 02 · Added context

Pages at this level last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page