# Level 03 · Workflows

_Software organizes the steps_

Connect model calls through predefined steps, branches, checks, and retries. A model can classify an input or evaluate a result to route the workflow; software still defines the available paths.


## Who decides the next step

Software defines the steps and allowed branches; model outputs can select among them.


## What is at this level

- [Prompt chaining](/gradient_ascent/techniques/prompt-chaining/) (sourced): Splitting a task into steps, each with its own prompt.
- [Routing](/gradient_ascent/techniques/routing/) (sourced): Sorting inputs and sending each one to the right prompt.
- [Parallel calls](/gradient_ascent/techniques/parallelization/) (sourced): Running several prompts at once and combining the results.
- [Write and check](/gradient_ascent/techniques/evaluator-optimizer/) (sourced): One prompt writes, another checks, and the loop repeats until the check passes.
- [Workflow graphs](/gradient_ascent/techniques/workflow-graphs/) (sourced): Describing a workflow as steps and the connections between them.
- [Human approval](/gradient_ascent/techniques/human-in-the-loop/) (sourced): Pausing for a person to approve or correct.

## Upgrade conditions

- **Prompt chaining → Single agent:** The steps cannot be written down in advance.
- **Routing → Function calling:** The right destination depends on something the message does not say, so the choice needs a lookup before it can be made.
- **Parallel calls → Lead agent and workers:** The subtasks cannot be known until the task is read.
- **Write and check → Review and debate:** One reviewer shares the author's blind spots.
- **Workflow graphs → Agent graphs:** A node needs its own loop and tools, not one call.

## Named products, tools and models


### Products

- Claude Code subagents — Anthropic · multi-agent feature of a coding agent
- Cline — Cline · open-source coding agent
- Dify — LangGenius · visual workflow builder
- Flowise — Flowise · visual workflow builder
- Grok Heavy — SpaceXAI · several agents answering one question
- Gumloop — Gumloop · visual workflow builder
- Jules — Google · coding agent
- Langflow — IBM · visual workflow builder
- Make — Celonis · automation service
- n8n — n8n · automation service, self-hostable
- Power Automate — Microsoft · automation service
- Replit Agent — Replit · coding agent
- Zapier — Zapier · automation service

### Tools

- Apache Airflow — open source · workflow engine
- Batch API — OpenAI · asynchronous batch processing api
- Batch API — Google · asynchronous batch processing api
- Deep Agents — LangChain · agent harness
- DSPy — Stanford NLP · prompt programs and optimizers
- Haystack — deepset · retrieval framework
- Inngest — Inngest · durable workflow engine
- LangChain — LangChain · application framework
- LangGraph — LangChain · graph framework
- Message Batches API — Anthropic · asynchronous batch processing api
- Prefect — Prefect · workflow engine
- Semantic Router — Aurelio Labs · library that routes a request by meaning
- Temporal — Temporal · durable workflow engine

### Models

- Jev — TypeSafe AI · system one decision model

## What is still unsolved at this level

_As of 09/19/2026. This block ages faster than the rest of the page._


### Prompt chaining

A wrong number introduced early in a fixed chain does not simply persist. It becomes a wrong calculation, then a sentence of prose, then a conclusion, and it gets harder to spot at every handoff, because each stage makes it look more like ordinary output.

**What people are trying:** Putting a check between stages rather than only at the end. One measurement through a four-stage pipeline found an early gate catches much more than the same check run once on the final result.

- [The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines](https://arxiv.org/abs/2608.14588) · arXiv · read 09/19/2026: "Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable and severe consequences."
  - The paper says multi-agent. The pipeline it measures is a fixed sequence of stages where no stage chooses what happens next, which is what this site calls prompt chaining.

### Routing

A router sorts a request into a branch before anything else runs, and routers still send plenty of requests somewhere other than where they would have been answered best. Under one benchmark that compared many of them on the same footing, several recent methods, commercial ones included, did not clearly beat a simple baseline.

**What people are trying:** Shared testbeds spanning many models and many routing methods, so routers are compared on the same task set instead of each paper's own. Some work scores cost and latency alongside accuracy rather than accuracy alone.

- [LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing](https://arxiv.org/abs/2601.07206) · arXiv · read 09/19/2026: "a substantial gap remains to the Oracle, driven primarily by persistent model-recall failures"

### Write and check

The check half of a write-and-check loop is usually another model call, and a model grading text written in its own style can favor it for reasons that have nothing to do with quality. Picking a more capable judge does not reliably fix it.

**What people are trying:** Measuring the bias directly, by grading pairs of answers built to differ in style and not in quality, so any preference shown is the bias. Structured multi-part rubrics are being tested as a way to cut it down.

- [Quantifying and Mitigating Self-Preference Bias of LLM Judges](https://arxiv.org/abs/2604.22891) · arXiv · read 09/19/2026: "Empirical analysis across 20 mainstream LLMs reveals that advanced capabilities are often uncorrelated, or even negatively correlated, with low SPB."

### Human approval

Putting a person in the loop is the easy part. How much they should be asked to decide, and how to keep them from waving through whatever the system already proposed, is not settled, and NIST says so in its own guidance rather than prescribing a split.

**What people are trying:** NIST's playbook asks teams to design the human role deliberately, account for the biases that make a reviewer agree too readily, and measure whether oversight is happening at all. Override rates and time spent per review are what tell you, rather than the presence of a checkpoint.

- [Map (AI RMF Playbook, AIRC)](https://airc.nist.gov/airmf-resources/playbook/map/) · NIST, Trustworthy and Responsible AI Resource Center · read 09/19/2026: "Questions remain about how to configure humans and automation for managing AI risks."
