Explore a candidate level for your workflow
About two minutes. Nothing you type here is sent anywhere; the whole thing runs in your browser. Answer for the workflow you want, including what should happen automatically. The result classifies one candidate design; it does not establish that this is the best solution.
Answer a few questions about the task
Each question classifies a candidate or explores the next level. Four short questions after that change the advice, not the level.
Can the job be done by a fixed rule, a search, or a plain form, with no model involved at all?
Consider fixed logic only if the complete workflow meets your automation and hands-on effort requirements. Technical feasibility alone does not establish the best fit.
JavaScript is off (or still loading), so here is the whole worksheet at once: every question, what each answer points to, and all eight possible results below it. With JavaScript on, this becomes one question at a time with a result built from your actual answers.
Can the job be done by a fixed rule, a search, or a plain form, with no model involved at all?
Consider fixed logic only if the complete workflow meets your automation and hands-on effort requirements. Technical feasibility alone does not establish the best fit.
- Yes. A rule, a lookup table, a search, or a plain form covers every case. — settles at Level 00 · Conventional software.
- No. It needs to understand or produce language, or use judgment a rule cannot capture. — not enough on its own: continue to Does one request, written well, with everything the model needs already in it, reliably get a good answer?
Does one request, written well, with everything the model needs already in it, reliably get a good answer?
This covers rewriting, drafting, or answering a question when all the facts are already in your message.
- Yes. One clear message, sent once, does it. — settles at Level 01 · Direct prompting.
- No. It needs information the model was not given: documents, current data, or something specific to us. — not enough on its own: continue to Does finding the right information and putting it in front of the model answer this, where one search or one set of documents is enough?
Does finding the right information and putting it in front of the model answer this, where one search or one set of documents is enough?
This covers answering from a manual, a knowledge base, or a set of notes, where a single lookup finds what is needed.
- Yes. Hand it the right document or search result and it answers well. — settles at Level 02 · Added context.
- No. It takes more than one step, or the steps have to happen in a set order. — not enough on its own: continue to Can you define the steps and allowed branches in advance, including how model outputs choose among those branches?
Can you define the steps and allowed branches in advance, including how model outputs choose among those branches?
This covers sorting inputs into categories, chaining prompts, and checking drafts. A model can select a predefined branch; software owns the available paths and stopping rules.
- Yes. Software defines the steps and branches, even if a model helps choose a branch. — settles at Level 03 · Workflows.
- No. It has to act on the world while it works: look something up live, run code, or operate something. — not enough on its own: continue to Can the model do this by picking one bounded action (calling a function, running a lookup, clicking one thing), with your code carrying out that action and stopping there?
Can the model do this by picking one bounded action (calling a function, running a lookup, clicking one thing), with your code carrying out that action and stopping there?
The model chooses which action to take and with what details, but only one action, and your code performs it and hands back the result.
- Yes. One action; your code runs it and the task is done. — settles at Level 04 · Tool use.
- No. What happens next depends on what the last action returned, and the model has to decide that for itself, more than once. — not enough on its own: continue to Can the steps NOT be known in advance, so the model has to plan, act, look at the result, and decide for itself what to do next and when it is finished?
Can the steps NOT be known in advance, so the model has to plan, act, look at the result, and decide for itself what to do next and when it is finished?
This is the difference between following a plan you wrote and working one out as it goes, the way debugging or open-ended research does.
- Yes. It plans, acts, checks its own result, and decides for itself when it is done. — settles at Level 05 · Agent loops.
- No. Even one model working alone in a loop is not enough for this. — not enough on its own: continue to Does the work need to be split across more than one agent or checked independently, or does it have to start on its own or keep running for a long time?
Does the work need to be split across more than one agent or checked independently, or does it have to start on its own or keep running for a long time?
Pick the closest answer. "Independent" means an agent with its own access, not a second pass by the same one.
- It is too big for one agent to hold, or the subtasks cannot be known until the work is split among several agents. — settles at Level 06 · Teams of Agents.
- It needs a second, independent agent to check the work, with its own access: one agent checking itself shares its own blind spots. — settles at Level 06 · Teams of Agents.
- It has to start without being asked, run on its own schedule, or keep going for days or longer. — settles at Level 07 · Always-on agents.
If the model gets this wrong, what does that cost?
Does not change the level. Answered after the level is settled, and only changes the cautions shown with the result.
- Not much. Someone notices quickly and it is easy to fix.
- Real rework, a bad customer moment, or a delay before someone catches it. — A wrong answer here costs real time or trust. Check the work before it goes out, and keep a record of what was decided and why.
- Money, legal exposure, safety, or something that cannot be undone. — A wrong answer here is expensive or dangerous. Keep a person in the loop before anything ships, and be able to say why the model was trusted with it.
How easily can someone check the result before it is used?
Does not change the level. Answered after the level is settled, and only changes the cautions shown with the result.
- Easily and fast: a glance, or a simple test, confirms it.
- It takes real effort: reading closely, or running it, to know if it is right. — Build the check into the process instead of relying on a read-through. A standing eval set catches drift a one-off spot check misses.
- It is hard or impossible to check: there is no ground truth, or checking takes as long as doing the job. — When nobody can easily confirm the result, treat it as unverified until proven otherwise, and keep a person responsible for what happens with it.
If the action turns out to be wrong, can it be undone?
Does not change the level. Answered after the level is settled, and only changes the cautions shown with the result.
- Yes, easily: nothing has shipped, been spent, or changed yet.
- With some effort: it can be corrected, but it takes real work. — Log every action taken so a wrong one can be traced and reversed without guessing what happened.
- No: money moved, a message went out, or something was deleted. — An action that cannot be undone needs approval before it happens, not review after the fact.
Does private or sensitive data leave your own machine to do this?
Does not change the level. Answered after the level is settled, and only changes the cautions shown with the result.
- No: everything runs on infrastructure you control.
- Yes: a cloud model or a third-party service sees it. — Know what leaves the machine and who can see it. Check the provider's data-handling terms before sending anything sensitive.
Conventional software
Use ordinary code, search, forms, or a task-specific statistical model when they solve the problem. No generative model is required; classical machine learning can belong here too.
Keep the household paperwork straight
Organize renewal dates, file names, category totals, and reminders with ordinary code. No model is needed; extracting information from scanned bills is a separate task.
Check measurements against limits, and chart what drifts
Use code to calculate limits, yield, process capability, and trends across lots and fixtures. The pass/fail decision stays deterministic; no model is involved.
Sweep a design over its corners and report the margins
Sweep prototype boards across line, load, and temperature. Code calculates margins, uncertainty, and guardbanded verdicts; no model is needed.
Direct prompting
Give the model instructions and receive a response. A conversation repeats this interaction, with a person directing each turn. Prompting, structured output, reasoning, and multimodal inputs can all fit this pattern.
Images, audio and video
SourcedGiving the model images, audio, video and documents, and getting them back.
Turn a meeting transcript into decisions and owners
Turn a transcript into decisions, owners, and open questions in one model call. Someone who attended reviews the draft before it is shared.
Watch a topic for new work and summarize what turns up
Code detects new records from fixed sources. One model call summarizes each new title and abstract; code attaches the original citation. It does not follow references or choose new searches.
Assemble a weekly status report from several systems
Code assembles the weekly figures; one model call drafts the report. Checks flag unsupported numbers and missing required facts, then a person reviews and sends it.
Turn a measurement session into a report somebody can review
Turn computed measurements and notebook notes into a report. Code owns the figures, the model writes the prose, and a person checks the finished draft.
Added context
Add relevant documents, retrieved passages, or stored information to the current request. This supplies context beyond the model’s training without retraining it. Missing, stale, or misleading material can still produce a poor answer.
Retrieval-augmented generation (RAG)
MeasuredSearching your documents and giving the results to the model.
Knowledge graphs and GraphRAG
SourcedStoring facts as entities and relations, for questions that span several documents.
Answer questions about a set of documents
Uses RAG, structured output and an eval set. Level 2 is enough because a single search answers most questions.
Answer questions from a datasheet, a test spec and a change notice
Retrieval over the documents an engineer already has, answered with citations that can be checked. The case that matters is a change notice contradicting the datasheet on one number, where the right answer depends on the board revision. Level 2 is enough because one search finds the passage.
Workflows
Connect model calls through predefined steps, branches, checks, and retries. A model can classify an input or evaluate a result to route the workflow; software still defines the available paths.
Write and check
SourcedOne prompt writes, another checks, and the loop repeats until the check passes.
Sort an inbox
Sorts mail into fixed categories and produces structured output. A person approves anything that gets sent. The categories are known in advance, so an agent is not needed.
Turn photos and PDFs into records
Reads the image or PDF, fills a fixed schema, and saves the record once a person confirms it.
Voice notes into structured entries
Transcribes a voice note, splits it into steps, and turns each step into a structured entry that an eval set checks for accuracy.
Nightly source monitor
Runs on a timer, diffs a set of public pages in code, and asks a model one question about each change. Level 3: the schedule and the checkpoint are infrastructure, not agency.
Drafting with a reviewer
One prompt drafts a piece of writing and another checks it against a rubric, repeating until the draft passes.
Match invoices to purchase orders
Extract invoice fields, then use code to match purchase orders and compare amounts. Differences go to a person; the model never decides whether the totals reconcile.
Check an agreement against your own checklist
Check an agreement against a fixed checklist, with cited clauses for each finding. Merge the findings for a person to review.
Turn an incident write-up into a runbook
Turn an incident write-up into a timeline and repeatable steps. Check owners and success criteria, then ask the incident lead to approve it.
Turn a script into a shot list
Split a script into scenes and shots, then check that every line is covered and every shot has a source. A person reviews the plan; drawing frames is a separate task.
Keep a tracker document current from several sources
Keep a shared tracker current through source comparisons and a review queue. Model proposals and changes to human-written fields need approval; missing evidence is flagged.
Pull an instrument's accuracy table out of its manual
Extract specification rows from a manual, validate their structure, and calculate uncertainty in code. A person verifies ranges, intervals, and conditions against the source.
Sort failing units and operator notes into causes
Failing measurements and free-text operator notes are sorted into the causes the failure analysis guide already lists, then routed. A person confirms before anything is scrapped or reworked. The categories are known in advance, so this is classification into fixed classes and not an agent.
Check a board against the design rules document
A bill of materials and a netlist summary are checked rule by rule against the written design rules. One pass drafts findings, a second checks each finding against the rule text it cites and drops the ones that cite nothing. Level 3, because code decides every step and the rules do not change between boards. This is the rule check that happens before a review meeting, not the design review report itself: for the report, and the characterization data behind it, see the two recipes this page links in its first paragraph.
Turn a requirements list into a test plan
A fixed chain: read the requirements, propose a test for each, build the traceability table, then check that every requirement has a test and every test names a requirement. A person approves before any of it is adopted. The order of the steps is known in advance, which is what keeps this at level 3.
Tool use
The model can request a search, calculation, code execution, or an action in another application. Software enforces permissions, performs the action, and returns the result. Tool use alone does not create an ongoing agent loop.
Support desk
Routes an incoming ticket, searches the documentation for an answer, calls a tool when an action is needed, and hands off to a person when it is unsure.
Plain-language maintenance log
Turns a plain-language description of work done into a structured log entry, saved with a tool call and linked to the equipment it concerns through a small knowledge graph.
Draft an instrument control script from its programming manual
The model drafts commands from the manual for that instrument; code checks every one against the documented command set, runs the script on the simulated instrument, and feeds the errors back for another pass. A person bench-checks before it drives real hardware, and every set point goes through a code-side envelope.
Ask questions of a production test log
Starts where the dashboard stopped: limits, yield and Cpk are already charted and did not answer the question. The model writes analysis code that runs in a sandbox over the CSV, and a person reads the code as well as the answer. Includes the trap of a column in millivolts under a header that says volts.
Agent loops
The model uses the goal and observed results to choose an action, revise its approach, or finish. Software executes tools and enforces permissions, approvals, and stopping limits. A run can stop because it is complete, blocked, or out of budget.
The agent harness
SourcedEverything around the model in an agent: the loop, tools, context handling, permissions, caps and sandbox.
Write a research brief with citations
Uses agentic RAG to find sources and a fixed check on every claim against the section it cites. It needs level 5 for the searching; the checking is level 3.
Coding assistant on your own repo
A coding agent that reads, edits, runs and tests code in your repository, using skills for repeated tasks and a safety review before anything ships.
Data analysis by conversation
A single agent writes and runs code against a dataset, one question at a time, to answer questions a fixed query could not anticipate.
Plan a trip and hold the bookings
Checking what is available, what is open and what connects takes a different number of steps every time, which is what level 5 is for. Read-only lookups run unattended; anything that spends money stops for a person, with the price and the cancellation terms in front of them.
Work a bring-up problem at the bench
An agent with read-only tools, instrument queries, the test log and the datasheet, works a low output down to a cause and proposes the next measurement. Queries run unattended; anything that sets a voltage, a current limit or an output goes through the envelope and a person. Level 5 because each measurement depends on the last.
Teams of Agents
Agents divide, coordinate, or review work across separate contexts. A coordinator can combine their findings, and the agents may use the same model or different models. Coordination adds overhead, and separate reviewers can still make correlated mistakes.
Grade against a rubric, with a second reader
Two independent reviewers apply the same rubric. Disagreements go to the teacher rather than being averaged away.
Always-on agents
Saved state, schedules, and events let an agent start or resume work without a fresh chat message each time. The model need not run continuously, and a dedicated computer or agent team is optional. Permissions, human approvals, monitoring, and stop controls still apply.
A team of personal assistants
Several always-on agents split personal tasks among themselves, sharing memory and staying inside the same safety rules.