Worksheet

Explore a candidate level for your workflow

About two minutes. Nothing you type here is sent anywhere; the whole thing runs in your browser. Answer for the workflow you want, including what should happen automatically. The result classifies one candidate design; it does not establish that this is the best solution.

Two minutes, no code

Answer a few questions about the task

Each question classifies a candidate or explores the next level. Four short questions after that change the advice, not the level.

Question 1 of up to 11

Can the job be done by a fixed rule, a search, or a plain form, with no model involved at all?

Consider fixed logic only if the complete workflow meets your automation and hands-on effort requirements. Technical feasibility alone does not establish the best fit.

JavaScript is off (or still loading), so here is the whole worksheet at once: every question, what each answer points to, and all eight possible results below it. With JavaScript on, this becomes one question at a time with a result built from your actual answers.

  1. Can the job be done by a fixed rule, a search, or a plain form, with no model involved at all?

    Consider fixed logic only if the complete workflow meets your automation and hands-on effort requirements. Technical feasibility alone does not establish the best fit.

  2. Does one request, written well, with everything the model needs already in it, reliably get a good answer?

    This covers rewriting, drafting, or answering a question when all the facts are already in your message.

  3. Does finding the right information and putting it in front of the model answer this, where one search or one set of documents is enough?

    This covers answering from a manual, a knowledge base, or a set of notes, where a single lookup finds what is needed.

  4. Can you define the steps and allowed branches in advance, including how model outputs choose among those branches?

    This covers sorting inputs into categories, chaining prompts, and checking drafts. A model can select a predefined branch; software owns the available paths and stopping rules.

  5. Can the model do this by picking one bounded action (calling a function, running a lookup, clicking one thing), with your code carrying out that action and stopping there?

    The model chooses which action to take and with what details, but only one action, and your code performs it and hands back the result.

  6. Can the steps NOT be known in advance, so the model has to plan, act, look at the result, and decide for itself what to do next and when it is finished?

    This is the difference between following a plan you wrote and working one out as it goes, the way debugging or open-ended research does.

  7. Does the work need to be split across more than one agent or checked independently, or does it have to start on its own or keep running for a long time?

    Pick the closest answer. "Independent" means an agent with its own access, not a second pass by the same one.

    • It is too big for one agent to hold, or the subtasks cannot be known until the work is split among several agents. — settles at Level 06 · Teams of Agents.
    • It needs a second, independent agent to check the work, with its own access: one agent checking itself shares its own blind spots. — settles at Level 06 · Teams of Agents.
    • It has to start without being asked, run on its own schedule, or keep going for days or longer. — settles at Level 07 · Always-on agents.
  8. If the model gets this wrong, what does that cost?

    Does not change the level. Answered after the level is settled, and only changes the cautions shown with the result.

    • Not much. Someone notices quickly and it is easy to fix.
    • Real rework, a bad customer moment, or a delay before someone catches it. — A wrong answer here costs real time or trust. Check the work before it goes out, and keep a record of what was decided and why.
    • Money, legal exposure, safety, or something that cannot be undone. — A wrong answer here is expensive or dangerous. Keep a person in the loop before anything ships, and be able to say why the model was trusted with it.
  9. How easily can someone check the result before it is used?

    Does not change the level. Answered after the level is settled, and only changes the cautions shown with the result.

    • Easily and fast: a glance, or a simple test, confirms it.
    • It takes real effort: reading closely, or running it, to know if it is right. — Build the check into the process instead of relying on a read-through. A standing eval set catches drift a one-off spot check misses.
    • It is hard or impossible to check: there is no ground truth, or checking takes as long as doing the job. — When nobody can easily confirm the result, treat it as unverified until proven otherwise, and keep a person responsible for what happens with it.
  10. If the action turns out to be wrong, can it be undone?

    Does not change the level. Answered after the level is settled, and only changes the cautions shown with the result.

    • Yes, easily: nothing has shipped, been spent, or changed yet.
    • With some effort: it can be corrected, but it takes real work. — Log every action taken so a wrong one can be traced and reversed without guessing what happened.
    • No: money moved, a message went out, or something was deleted. — An action that cannot be undone needs approval before it happens, not review after the fact.
  11. Does private or sensitive data leave your own machine to do this?

    Does not change the level. Answered after the level is settled, and only changes the cautions shown with the result.

    • No: everything runs on infrastructure you control.
    • Yes: a cloud model or a third-party service sees it. — Know what leaves the machine and who can see it. Check the provider's data-handling terms before sending anything sensitive.
The eight resultsone per level
Level 00

Conventional software

Use ordinary code, search, forms, or a task-specific statistical model when they solve the problem. No generative model is required; classical machine learning can belong here too.

Who decides the next stepYour software applies rules, lookups, or established algorithms.
Why not level 01, Direct prompting?Give the model instructions and receive a response. A conversation repeats this interaction, with a person directing each turn. Prompting, structured output, reasoning, and multimodal inputs can all fit this pattern. It would take more to build, more to test and more ways to fail quietly — worth it only once level 00 has actually fallen short, not because it might. A proposal that skips several levels has to clear that same test at every level in between.
Techniques at this level1

When not to use a model

Sourced

How to tell when ordinary code, search or a form is enough.

Recipes that top out here3
Level 0

Keep the household paperwork straight

Organize renewal dates, file names, category totals, and reminders with ordinary code. No model is needed; extracting information from scanned bills is a separate task.

Needs level 0
↗
Level 0

Check measurements against limits, and chart what drifts

Use code to calculate limits, yield, process capability, and trends across lots and fixtures. The pass/fail decision stays deterministic; no model is involved.

Needs level 0
↗
Level 0

Sweep a design over its corners and report the margins

Sweep prototype boards across line, load, and temperature. Code calculates margins, uncertainty, and guardbanded verdicts; no model is needed.

Needs level 0
↗
Level 01

Direct prompting

Give the model instructions and receive a response. A conversation repeats this interaction, with a person directing each turn. Prompting, structured output, reasoning, and multimodal inputs can all fit this pattern.

Who decides the next stepYou choose the request; the model generates a response.
Why not level 02, Added context?Add relevant documents, retrieved passages, or stored information to the current request. This supplies context beyond the model’s training without retraining it. Missing, stale, or misleading material can still produce a poor answer. It would take more to build, more to test and more ways to fail quietly — worth it only once level 01 has actually fallen short, not because it might. A proposal that skips several levels has to clear that same test at every level in between.
Techniques at this level5

Chat

Sourced

Asking a model a question in a chat app.

Prompt engineering

Sourced

Writing instructions that get consistent results.

Structured output

Sourced

Getting answers in a fixed format such as JSON.

Reasoning at answer time

Sourced

Letting the model think for longer before it answers.

Images, audio and video

Sourced

Giving the model images, audio, video and documents, and getting them back.

Recipes that top out here4
Level 1

Turn a meeting transcript into decisions and owners

Turn a transcript into decisions, owners, and open questions in one model call. Someone who attended reviews the draft before it is shared.

Needs level 1
↗
Level 0 + Level 1

Watch a topic for new work and summarize what turns up

Code detects new records from fixed sources. One model call summarizes each new title and abstract; code attaches the original citation. It does not follow references or choose new searches.

Needs level 1
↗
Level 0 + Level 1

Assemble a weekly status report from several systems

Code assembles the weekly figures; one model call drafts the report. Checks flag unsupported numbers and missing required facts, then a person reviews and sends it.

Needs level 1
↗
Level 0 + Level 1

Turn a measurement session into a report somebody can review

Turn computed measurements and notebook notes into a report. Code owns the figures, the model writes the prose, and a person checks the finished draft.

Needs level 1
↗
Level 02

Added context

Add relevant documents, retrieved passages, or stored information to the current request. This supplies context beyond the model’s training without retraining it. Missing, stale, or misleading material can still produce a poor answer.

Who decides the next stepYou or your software select the information supplied to the model.
Why not level 03, Workflows?Connect model calls through predefined steps, branches, checks, and retries. A model can classify an input or evaluate a result to route the workflow; software still defines the available paths. It would take more to build, more to test and more ways to fail quietly — worth it only once level 02 has actually fallen short, not because it might. A proposal that skips several levels has to clear that same test at every level in between.
Techniques at this level5

Context engineering

Sourced

Deciding what goes into the request, and caching the parts that repeat.

Embeddings and search

Sourced

Finding text by meaning instead of by keyword.

Retrieval-augmented generation (RAG)

Measured

Searching your documents and giving the results to the model.

Knowledge graphs and GraphRAG

Sourced

Storing facts as entities and relations, for questions that span several documents.

Memory

Sourced

Keeping information from one conversation to the next.

Recipes that top out here2
Level 1 + Level 2

Answer questions about a set of documents

Uses RAG, structured output and an eval set. Level 2 is enough because a single search answers most questions.

Needs level 2
↗
Level 1 + Level 2

Answer questions from a datasheet, a test spec and a change notice

Retrieval over the documents an engineer already has, answered with citations that can be checked. The case that matters is a change notice contradicting the datasheet on one number, where the right answer depends on the board revision. Level 2 is enough because one search finds the passage.

Needs level 2
↗
Level 03

Workflows

Connect model calls through predefined steps, branches, checks, and retries. A model can classify an input or evaluate a result to route the workflow; software still defines the available paths.

Who decides the next stepSoftware defines the steps and allowed branches; model outputs can select among them.
Why not level 04, Tool use?The model can request a search, calculation, code execution, or an action in another application. Software enforces permissions, performs the action, and returns the result. Tool use alone does not create an ongoing agent loop. It would take more to build, more to test and more ways to fail quietly — worth it only once level 03 has actually fallen short, not because it might. A proposal that skips several levels has to clear that same test at every level in between.
Techniques at this level6

Prompt chaining

Sourced

Splitting a task into steps, each with its own prompt.

Routing

Sourced

Sorting inputs and sending each one to the right prompt.

Parallel calls

Sourced

Running several prompts at once and combining the results.

Write and check

Sourced

One prompt writes, another checks, and the loop repeats until the check passes.

Workflow graphs

Sourced

Describing a workflow as steps and the connections between them.

Human approval

Sourced

Pausing for a person to approve or correct.

Recipes that top out here14
Level 1 + Level 3

Sort an inbox

Sorts mail into fixed categories and produces structured output. A person approves anything that gets sent. The categories are known in advance, so an agent is not needed.

Needs level 3
↗
Level 1 + Level 3

Turn photos and PDFs into records

Reads the image or PDF, fills a fixed schema, and saves the record once a person confirms it.

Needs level 3
↗
Level 1 + Level 3

Voice notes into structured entries

Transcribes a voice note, splits it into steps, and turns each step into a structured entry that an eval set checks for accuracy.

Needs level 3
↗
Level 1 + Level 3

Nightly source monitor

Runs on a timer, diffs a set of public pages in code, and asks a model one question about each change. Level 3: the schedule and the checkpoint are infrastructure, not agency.

Needs level 3
↗
Level 3

Drafting with a reviewer

One prompt drafts a piece of writing and another checks it against a rubric, repeating until the draft passes.

Needs level 3
↗
Level 0 + Level 1 + Level 3

Match invoices to purchase orders

Extract invoice fields, then use code to match purchase orders and compare amounts. Differences go to a person; the model never decides whether the totals reconcile.

Needs level 3
↗
Level 1 + Level 3

Check an agreement against your own checklist

Check an agreement against a fixed checklist, with cited clauses for each finding. Merge the findings for a person to review.

Needs level 3
↗
Level 1 + Level 3

Turn an incident write-up into a runbook

Turn an incident write-up into a timeline and repeatable steps. Check owners and success criteria, then ask the incident lead to approve it.

Needs level 3
↗
Level 1 + Level 3

Turn a script into a shot list

Split a script into scenes and shots, then check that every line is covered and every shot has a source. A person reviews the plan; drawing frames is a separate task.

Needs level 3
↗
Level 0 + Level 1 + Level 3

Keep a tracker document current from several sources

Keep a shared tracker current through source comparisons and a review queue. Model proposals and changes to human-written fields need approval; missing evidence is flagged.

Needs level 3
↗
Level 0 + Level 1 + Level 3

Pull an instrument's accuracy table out of its manual

Extract specification rows from a manual, validate their structure, and calculate uncertainty in code. A person verifies ranges, intervals, and conditions against the source.

Needs level 3
↗
Level 1 + Level 3

Sort failing units and operator notes into causes

Failing measurements and free-text operator notes are sorted into the causes the failure analysis guide already lists, then routed. A person confirms before anything is scrapped or reworked. The categories are known in advance, so this is classification into fixed classes and not an agent.

Needs level 3
↗
Level 1 + Level 3

Check a board against the design rules document

A bill of materials and a netlist summary are checked rule by rule against the written design rules. One pass drafts findings, a second checks each finding against the rule text it cites and drops the ones that cite nothing. Level 3, because code decides every step and the rules do not change between boards. This is the rule check that happens before a review meeting, not the design review report itself: for the report, and the characterization data behind it, see the two recipes this page links in its first paragraph.

Needs level 3
↗
Level 1 + Level 3

Turn a requirements list into a test plan

A fixed chain: read the requirements, propose a test for each, build the traceability table, then check that every requirement has a test and every test names a requirement. A person approves before any of it is adopted. The order of the steps is known in advance, which is what keeps this at level 3.

Needs level 3
↗
Level 04

Tool use

The model can request a search, calculation, code execution, or an action in another application. Software enforces permissions, performs the action, and returns the result. Tool use alone does not create an ongoing agent loop.

Who decides the next stepThe model requests an action; software checks and executes it.
Why not level 05, Agent loops?The model uses the goal and observed results to choose an action, revise its approach, or finish. Software executes tools and enforces permissions, approvals, and stopping limits. A run can stop because it is complete, blocked, or out of budget. It would take more to build, more to test and more ways to fail quietly — worth it only once level 04 has actually fallen short, not because it might. A proposal that skips several levels has to clear that same test at every level in between.
Techniques at this level4

Function calling

Sourced

Letting the model call functions that you define.

Code execution

Sourced

Letting the model write code and run it in a sandbox.

Model Context Protocol

Sourced

A standard way to connect models to tools and data.

Computer and browser use

Sourced

Letting the model operate a screen, a mouse and a keyboard.

Recipes that top out here4
Level 2 + Level 3 + Level 4

Support desk

Routes an incoming ticket, searches the documentation for an answer, calls a tool when an action is needed, and hands off to a person when it is unsure.

Needs level 4
↗
Level 1 + Level 2 + Level 4

Plain-language maintenance log

Turns a plain-language description of work done into a structured log entry, saved with a tool call and linked to the equipment it concerns through a small knowledge graph.

Needs level 4
↗
Level 2 + Level 3 + Level 4

Draft an instrument control script from its programming manual

The model drafts commands from the manual for that instrument; code checks every one against the documented command set, runs the script on the simulated instrument, and feeds the errors back for another pass. A person bench-checks before it drives real hardware, and every set point goes through a code-side envelope.

Needs level 4
↗
Level 4

Ask questions of a production test log

Starts where the dashboard stopped: limits, yield and Cpk are already charted and did not answer the question. The model writes analysis code that runs in a sandbox over the CSV, and a person reads the code as well as the answer. Includes the trap of a column in millivolts under a header that says volts.

Needs level 4
↗
Level 05

Agent loops

The model uses the goal and observed results to choose an action, revise its approach, or finish. Software executes tools and enforces permissions, approvals, and stopping limits. A run can stop because it is complete, blocked, or out of budget.

Who decides the next stepThe model chooses the next step within limits enforced by software.
Why not level 06, Teams of Agents?Agents divide, coordinate, or review work across separate contexts. A coordinator can combine their findings, and the agents may use the same model or different models. Coordination adds overhead, and separate reviewers can still make correlated mistakes. It would take more to build, more to test and more ways to fail quietly — worth it only once level 05 has actually fallen short, not because it might. A proposal that skips several levels has to clear that same test at every level in between.
Techniques at this level6

Single agent

Sourced

A model that plans, acts and checks its own work in a loop.

The agent harness

Sourced

Everything around the model in an agent: the loop, tools, context handling, permissions, caps and sandbox.

Agentic RAG and deep research

Measured

An agent that runs its own searches until it has an answer.

Coding agents

Sourced

Agents that read, write, run and test code.

Skills

Sourced

Reusable instructions that an agent loads when it needs them.

Voice agents

Sourced

Agents you talk to in real time.

Recipes that top out here5
Level 3 + Level 5

Write a research brief with citations

Uses agentic RAG to find sources and a fixed check on every claim against the section it cites. It needs level 5 for the searching; the checking is level 3.

Needs level 5
↗
Level 5

Coding assistant on your own repo

A coding agent that reads, edits, runs and tests code in your repository, using skills for repeated tasks and a safety review before anything ships.

Needs level 5
↗
Level 4 + Level 5

Data analysis by conversation

A single agent writes and runs code against a dataset, one question at a time, to answer questions a fixed query could not anticipate.

Needs level 5
↗
Level 3 + Level 4 + Level 5

Plan a trip and hold the bookings

Checking what is available, what is open and what connects takes a different number of steps every time, which is what level 5 is for. Read-only lookups run unattended; anything that spends money stops for a person, with the price and the cancellation terms in front of them.

Needs level 5
↗
Level 3 + Level 4 + Level 5

Work a bring-up problem at the bench

An agent with read-only tools, instrument queries, the test log and the datasheet, works a low output down to a cause and proposes the next measurement. Queries run unattended; anything that sets a voltage, a current limit or an output goes through the envelope and a person. Level 5 because each measurement depends on the last.

Needs level 5
↗
Level 06

Teams of Agents

Agents divide, coordinate, or review work across separate contexts. A coordinator can combine their findings, and the agents may use the same model or different models. Coordination adds overhead, and separate reviewers can still make correlated mistakes.

Who decides the next stepSeveral agents coordinate, delegate, or review work; they can use the same underlying model.
Why not level 07, Always-on agents?Saved state, schedules, and events let an agent start or resume work without a fresh chat message each time. The model need not run continuously, and a dedicated computer or agent team is optional. Permissions, human approvals, monitoring, and stop controls still apply. It would take more to build, more to test and more ways to fail quietly — worth it only once level 06 has actually fallen short, not because it might. A proposal that skips several levels has to clear that same test at every level in between.
Techniques at this level3

Lead agent and workers

Sourced

A lead agent splits the task and hands parts to other agents.

Agent graphs

Sourced

Describing a team of agents and how work passes between them.

Review and debate

Sourced

Agents that check, or argue with, each other's work.

Recipes that top out here1
Level 1 + Level 3 + Level 6

Grade against a rubric, with a second reader

Two independent reviewers apply the same rubric. Disagreements go to the teacher rather than being averaged away.

Needs level 6
↗
Level 07

Always-on agents

Saved state, schedules, and events let an agent start or resume work without a fresh chat message each time. The model need not run continuously, and a dedicated computer or agent team is optional. Permissions, human approvals, monitoring, and stop controls still apply.

Who decides the next stepSoftware triggers and resumes runs; agents decide what to do within their standing instructions.
Techniques at this level4

Long-running tasks

Sourced

Tasks that run for hours or days.

Always-on assistants

Sourced

Agents that resume work across sessions, schedules, and events.

Organizations of agents

Sourced

Large groups of agents with roles and shared goals.

Robots and machines

Sourced

Models that control robots and other machines.

Recipes that top out here1
Level 2 + Level 6 + Level 7

A team of personal assistants

Several always-on agents split personal tasks among themselves, sharing memory and staying inside the same safety rules.

Needs level 7
↗