Common tasks and the levels they need
A recipe is a use case composed from techniques, with the reasoning for each choice. Every recipe names the techniques it uses.
For engineers: test, measurement, design and software
11 recipes, all Sourced, ordered by the level they need. They run on one bench, used three ways: production test, engineering test and precise measurement. The bench is an invented regulator board, four instruments with a SCPI command set, a test specification, a production run and a characterization sweep. The first two are a pair, one per setting, and neither uses a model at all, which is the honest starting point for most of this work.
Check measurements against limits, and chart what drifts
Use code to calculate limits, yield, process capability, and trends across lots and fixtures. The pass/fail decision stays deterministic; no model is involved.
Sweep a design over its corners and report the margins
Sweep prototype boards across line, load, and temperature. Code calculates margins, uncertainty, and guardbanded verdicts; no model is needed.
Turn a measurement session into a report somebody can review
Turn computed measurements and notebook notes into a report. Code owns the figures, the model writes the prose, and a person checks the finished draft.
Answer questions from a datasheet, a test spec and a change notice
Retrieval over the documents an engineer already has, answered with citations that can be checked. The case that matters is a change notice contradicting the datasheet on one number, where the right answer depends on the board revision. Level 2 is enough because one search finds the passage.
Pull an instrument's accuracy table out of its manual
Extract specification rows from a manual, validate their structure, and calculate uncertainty in code. A person verifies ranges, intervals, and conditions against the source.
Sort failing units and operator notes into causes
Failing measurements and free-text operator notes are sorted into the causes the failure analysis guide already lists, then routed. A person confirms before anything is scrapped or reworked. The categories are known in advance, so this is classification into fixed classes and not an agent.
Check a board against the design rules document
A bill of materials and a netlist summary are checked rule by rule against the written design rules. One pass drafts findings, a second checks each finding against the rule text it cites and drops the ones that cite nothing. Level 3, because code decides every step and the rules do not change between boards. This is the rule check that happens before a review meeting, not the design review report itself: for the report, and the characterization data behind it, see the two recipes this page links in its first paragraph.
Turn a requirements list into a test plan
A fixed chain: read the requirements, propose a test for each, build the traceability table, then check that every requirement has a test and every test names a requirement. A person approves before any of it is adopted. The order of the steps is known in advance, which is what keeps this at level 3.
Draft an instrument control script from its programming manual
The model drafts commands from the manual for that instrument; code checks every one against the documented command set, runs the script on the simulated instrument, and feeds the errors back for another pass. A person bench-checks before it drives real hardware, and every set point goes through a code-side envelope.
Ask questions of a production test log
Starts where the dashboard stopped: limits, yield and Cpk are already charted and did not answer the question. The model writes analysis code that runs in a sandbox over the CSV, and a person reads the code as well as the answer. Includes the trap of a column in millivolts under a header that says volts.
Work a bring-up problem at the bench
An agent with read-only tools, instrument queries, the test log and the datasheet, works a low output down to a cause and proposes the next measurement. Queries run unattended; anything that sets a voltage, a current limit or an output goes through the envelope and a person. Level 5 because each measurement depends on the last.
For everyone: documents, mail, notes and teams
23 recipes, all Sourced. The left column lists the levels a recipe draws on; the right names the highest one it needs.
Answer questions about a set of documents
Uses RAG, structured output and an eval set. Level 2 is enough because a single search answers most questions.
Sort an inbox
Sorts mail into fixed categories and produces structured output. A person approves anything that gets sent. The categories are known in advance, so an agent is not needed.
Write a research brief with citations
Uses agentic RAG to find sources and a fixed check on every claim against the section it cites. It needs level 5 for the searching; the checking is level 3.
Coding assistant on your own repo
A coding agent that reads, edits, runs and tests code in your repository, using skills for repeated tasks and a safety review before anything ships.
Turn photos and PDFs into records
Reads the image or PDF, fills a fixed schema, and saves the record once a person confirms it.
Voice notes into structured entries
Transcribes a voice note, splits it into steps, and turns each step into a structured entry that an eval set checks for accuracy.
Support desk
Routes an incoming ticket, searches the documentation for an answer, calls a tool when an action is needed, and hands off to a person when it is unsure.
Nightly source monitor
Runs on a timer, diffs a set of public pages in code, and asks a model one question about each change. Level 3: the schedule and the checkpoint are infrastructure, not agency.
Data analysis by conversation
A single agent writes and runs code against a dataset, one question at a time, to answer questions a fixed query could not anticipate.
Drafting with a reviewer
One prompt drafts a piece of writing and another checks it against a rubric, repeating until the draft passes.
A team of personal assistants
Several always-on agents split personal tasks among themselves, sharing memory and staying inside the same safety rules.
Plain-language maintenance log
Turns a plain-language description of work done into a structured log entry, saved with a tool call and linked to the equipment it concerns through a small knowledge graph.
Keep the household paperwork straight
Organize renewal dates, file names, category totals, and reminders with ordinary code. No model is needed; extracting information from scanned bills is a separate task.
Turn a meeting transcript into decisions and owners
Turn a transcript into decisions, owners, and open questions in one model call. Someone who attended reviews the draft before it is shared.
Match invoices to purchase orders
Extract invoice fields, then use code to match purchase orders and compare amounts. Differences go to a person; the model never decides whether the totals reconcile.
Check an agreement against your own checklist
Check an agreement against a fixed checklist, with cited clauses for each finding. Merge the findings for a person to review.
Turn an incident write-up into a runbook
Turn an incident write-up into a timeline and repeatable steps. Check owners and success criteria, then ask the incident lead to approve it.
Turn a script into a shot list
Split a script into scenes and shots, then check that every line is covered and every shot has a source. A person reviews the plan; drawing frames is a separate task.
Plan a trip and hold the bookings
Checking what is available, what is open and what connects takes a different number of steps every time, which is what level 5 is for. Read-only lookups run unattended; anything that spends money stops for a person, with the price and the cancellation terms in front of them.
Grade against a rubric, with a second reader
Two independent reviewers apply the same rubric. Disagreements go to the teacher rather than being averaged away.
Watch a topic for new work and summarize what turns up
Code detects new records from fixed sources. One model call summarizes each new title and abstract; code attaches the original citation. It does not follow references or choose new searches.
Assemble a weekly status report from several systems
Code assembles the weekly figures; one model call drafts the report. Checks flag unsupported numbers and missing required facts, then a person reviews and sends it.
Keep a tracker document current from several sources
Keep a shared tracker current through source comparisons and a review queue. Model proposals and changes to human-written fields need approval; missing evidence is flagged.
34 recipes across 24 techniques
Both domains in one grid. Every technique some recipe uses gets a column, grouped and colored by level; the 30 pages no recipe uses are left out. A filled cell means that recipe uses that technique. The grid scrolls sideways inside its own box; on a narrow screen it is a list instead.
| Recipe | Level 0 | Level 1 | Level 2 | Level 3 | Level 4 | Level 5 | Level 6 | Level 7 | Topics | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| When not to use a model | Prompt engineering | Structured output | Images, audio and video | Retrieval-augmented generation (RAG) | Knowledge graphs and GraphRAG | Memory | Prompt chaining | Routing | Parallel calls | Write and check | Human approval | Function calling | Code execution | Single agent | Agentic RAG and deep research | Coding agents | Skills | Agent graphs | Review and debate | Always-on assistants | Evals | Safety, privacy and governance | Operations | |
| Answer questions about a set of documents | uses Structured output | uses Retrieval-augmented generation (RAG) | uses Evals | |||||||||||||||||||||
| Sort an inbox | uses Structured output | uses Routing | uses Human approval | |||||||||||||||||||||
| Write a research brief with citations | uses Write and check | uses Agentic RAG and deep research | ||||||||||||||||||||||
| Coding assistant on your own repo | uses Coding agents | uses Skills | uses Safety, privacy and governance | |||||||||||||||||||||
| Turn photos and PDFs into records | uses Structured output | uses Images, audio and video | uses Human approval | |||||||||||||||||||||
| Voice notes into structured entries | uses Images, audio and video | uses Prompt chaining | uses Evals | |||||||||||||||||||||
| Support desk | uses Retrieval-augmented generation (RAG) | uses Routing | uses Human approval | uses Function calling | ||||||||||||||||||||
| Nightly source monitor | uses Structured output | uses Routing | uses Evals | uses Operations | ||||||||||||||||||||
| Data analysis by conversation | uses Code execution | uses Single agent | ||||||||||||||||||||||
| Drafting with a reviewer | uses Prompt chaining | uses Write and check | ||||||||||||||||||||||
| A team of personal assistants | uses Memory | uses Agent graphs | uses Always-on assistants | uses Safety, privacy and governance | ||||||||||||||||||||
| Plain-language maintenance log | uses Structured output | uses Knowledge graphs and GraphRAG | uses Function calling | |||||||||||||||||||||
| Keep the household paperwork straight | uses When not to use a model | |||||||||||||||||||||||
| Turn a meeting transcript into decisions and owners | uses Prompt engineering | uses Structured output | ||||||||||||||||||||||
| Match invoices to purchase orders | uses When not to use a model | uses Structured output | uses Human approval | |||||||||||||||||||||
| Check an agreement against your own checklist | uses Structured output | uses Parallel calls | uses Human approval | |||||||||||||||||||||
| Turn an incident write-up into a runbook | uses Structured output | uses Prompt chaining | uses Human approval | |||||||||||||||||||||
| Turn a script into a shot list | uses Structured output | uses Prompt chaining | ||||||||||||||||||||||
| Plan a trip and hold the bookings | uses Human approval | uses Function calling | uses Single agent | |||||||||||||||||||||
| Grade against a rubric, with a second reader | uses Structured output | uses Human approval | uses Review and debate | |||||||||||||||||||||
| Watch a topic for new work and summarize what turns up | uses When not to use a model | uses Prompt engineering | uses Structured output | uses Evals | uses Operations | |||||||||||||||||||
| Assemble a weekly status report from several systems | uses When not to use a model | uses Prompt engineering | uses Evals | uses Operations | ||||||||||||||||||||
| Keep a tracker document current from several sources | uses When not to use a model | uses Structured output | uses Human approval | uses Operations | ||||||||||||||||||||
| Check measurements against limits, and chart what drifts | uses When not to use a model | |||||||||||||||||||||||
| Sweep a design over its corners and report the margins | uses When not to use a model | |||||||||||||||||||||||
| Turn a measurement session into a report somebody can review | uses When not to use a model | uses Prompt engineering | ||||||||||||||||||||||
| Answer questions from a datasheet, a test spec and a change notice | uses Structured output | uses Retrieval-augmented generation (RAG) | ||||||||||||||||||||||
| Pull an instrument's accuracy table out of its manual | uses When not to use a model | uses Structured output | uses Human approval | |||||||||||||||||||||
| Sort failing units and operator notes into causes | uses Structured output | uses Routing | uses Human approval | |||||||||||||||||||||
| Check a board against the design rules document | uses Structured output | uses Write and check | ||||||||||||||||||||||
| Turn a requirements list into a test plan | uses Structured output | uses Prompt chaining | uses Human approval | |||||||||||||||||||||
| Draft an instrument control script from its programming manual | uses Retrieval-augmented generation (RAG) | uses Write and check | uses Code execution | |||||||||||||||||||||
| Ask questions of a production test log | uses Function calling | uses Code execution | ||||||||||||||||||||||
| Work a bring-up problem at the bench | uses Human approval | uses Function calling | uses Single agent | uses Safety, privacy and governance | ||||||||||||||||||||
| Recipes using it | 9 | 4 | 18 | 2 | 4 | 1 | 1 | 5 | 4 | 1 | 4 | 13 | 5 | 3 | 3 | 1 | 1 | 1 | 1 | 1 | 1 | 5 | 3 | 4 |
Answer questions about a set of documents
Sort an inbox
Write a research brief with citations
Coding assistant on your own repo
Turn photos and PDFs into records
Voice notes into structured entries
Support desk
Nightly source monitor
Data analysis by conversation
Drafting with a reviewer
A team of personal assistants
Plain-language maintenance log
Keep the household paperwork straight
Turn a meeting transcript into decisions and owners
Match invoices to purchase orders
Check an agreement against your own checklist
Turn an incident write-up into a runbook
Turn a script into a shot list
Plan a trip and hold the bookings
Grade against a rubric, with a second reader
Watch a topic for new work and summarize what turns up
Assemble a weekly status report from several systems
Keep a tracker document current from several sources
Check measurements against limits, and chart what drifts
Sweep a design over its corners and report the margins
Turn a measurement session into a report somebody can review
Answer questions from a datasheet, a test spec and a change notice
Pull an instrument's accuracy table out of its manual
Sort failing units and operator notes into causes
Check a board against the design rules document
Turn a requirements list into a test plan
Draft an instrument control script from its programming manual
Ask questions of a production test log
Work a bring-up problem at the bench
The columns say which techniques earn their place across many jobs and which are specialized. A recipe's own page gives the reasoning for each choice, and says why it does not need a higher level.