Recipes

Common tasks and the levels they need

A recipe is a use case composed from techniques, with the reasoning for each choice. Every recipe names the techniques it uses.

For engineers

For engineers: test, measurement, design and software

11 recipes, all Sourced, ordered by the level they need. They run on one bench, used three ways: production test, engineering test and precise measurement. The bench is an invented regulator board, four instruments with a SCPI command set, a test specification, a production run and a characterization sweep. The first two are a pair, one per setting, and neither uses a model at all, which is the honest starting point for most of this work.

Level 0

Check measurements against limits, and chart what drifts

Use code to calculate limits, yield, process capability, and trends across lots and fixtures. The pass/fail decision stays deterministic; no model is involved.

This example uses level 0
↗
Level 0

Sweep a design over its corners and report the margins

Sweep prototype boards across line, load, and temperature. Code calculates margins, uncertainty, and guardbanded verdicts; no model is needed.

This example uses level 0
↗
Level 0 + Level 1

Turn a measurement session into a report somebody can review

Turn computed measurements and notebook notes into a report. Code owns the figures, the model writes the prose, and a person checks the finished draft.

This example uses level 1
↗
Level 1 + Level 2

Answer questions from a datasheet, a test spec and a change notice

Retrieval over the documents an engineer already has, answered with citations that can be checked. The case that matters is a change notice contradicting the datasheet on one number, where the right answer depends on the board revision. Level 2 is enough because one search finds the passage.

This example uses level 2
↗
Level 0 + Level 1 + Level 3

Pull an instrument's accuracy table out of its manual

Extract specification rows from a manual, validate their structure, and calculate uncertainty in code. A person verifies ranges, intervals, and conditions against the source.

This example uses level 3
↗
Level 1 + Level 3

Sort failing units and operator notes into causes

Failing measurements and free-text operator notes are sorted into the causes the failure analysis guide already lists, then routed. A person confirms before anything is scrapped or reworked. The categories are known in advance, so this is classification into fixed classes and not an agent.

This example uses level 3
↗
Level 1 + Level 3

Check a board against the design rules document

A bill of materials and a netlist summary are checked rule by rule against the written design rules. One pass drafts findings, a second checks each finding against the rule text it cites and drops the ones that cite nothing. Level 3, because code decides every step and the rules do not change between boards. This is the rule check that happens before a review meeting, not the design review report itself: for the report, and the characterization data behind it, see the two recipes this page links in its first paragraph.

This example uses level 3
↗
Level 1 + Level 3

Turn a requirements list into a test plan

A fixed chain: read the requirements, propose a test for each, build the traceability table, then check that every requirement has a test and every test names a requirement. A person approves before any of it is adopted. The order of the steps is known in advance, which is what keeps this at level 3.

This example uses level 3
↗
Level 2 + Level 3 + Level 4

Draft an instrument control script from its programming manual

The model drafts commands from the manual for that instrument; code checks every one against the documented command set, runs the script on the simulated instrument, and feeds the errors back for another pass. A person bench-checks before it drives real hardware, and every set point goes through a code-side envelope.

This example uses level 4
↗
Level 4

Ask questions of a production test log

Starts where the dashboard stopped: limits, yield and Cpk are already charted and did not answer the question. The model writes analysis code that runs in a sandbox over the CSV, and a person reads the code as well as the answer. Includes the trap of a column in millivolts under a header that says volts.

This example uses level 4
↗
Level 3 + Level 4 + Level 5

Work a bring-up problem at the bench

An agent with read-only tools, instrument queries, the test log and the datasheet, works a low output down to a cause and proposes the next measurement. Queries run unattended; anything that sets a voltage, a current limit or an output goes through the envelope and a person. Level 5 because each measurement depends on the last.

This example uses level 5
↗
Everyday jobs

For everyone: documents, mail, notes and teams

23 recipes, all Sourced. The left column lists the levels a recipe draws on; the right names the highest one it needs.

Level 1 + Level 2

Answer questions about a set of documents

Uses RAG, structured output and an eval set. Level 2 is enough because a single search answers most questions.

Used as the running example
↗
Level 1 + Level 3

Sort an inbox

Sorts mail into fixed categories and produces structured output. A person approves anything that gets sent. The categories are known in advance, so an agent is not needed.

This example uses level 3
↗
Level 3 + Level 5

Write a research brief with citations

Uses agentic RAG to find sources and a fixed check on every claim against the section it cites. It needs level 5 for the searching; the checking is level 3.

This example uses level 5
↗
Level 5

Coding assistant on your own repo

A coding agent that reads, edits, runs and tests code in your repository, using skills for repeated tasks and a safety review before anything ships.

This example uses level 5
↗
Level 1 + Level 3

Turn photos and PDFs into records

Reads the image or PDF, fills a fixed schema, and saves the record once a person confirms it.

This example uses level 3
↗
Level 1 + Level 3

Voice notes into structured entries

Transcribes a voice note, splits it into steps, and turns each step into a structured entry that an eval set checks for accuracy.

This example uses level 3
↗
Level 2 + Level 3 + Level 4

Support desk

Routes an incoming ticket, searches the documentation for an answer, calls a tool when an action is needed, and hands off to a person when it is unsure.

This example uses level 4
↗
Level 1 + Level 3

Nightly source monitor

Runs on a timer, diffs a set of public pages in code, and asks a model one question about each change. Level 3: the schedule and the checkpoint are infrastructure, not agency.

This example uses level 3
↗
Level 4 + Level 5

Data analysis by conversation

A single agent writes and runs code against a dataset, one question at a time, to answer questions a fixed query could not anticipate.

This example uses level 5
↗
Level 3

Drafting with a reviewer

One prompt drafts a piece of writing and another checks it against a rubric, repeating until the draft passes.

This example uses level 3
↗
Level 2 + Level 6 + Level 7

A team of personal assistants

Several always-on agents split personal tasks among themselves, sharing memory and staying inside the same safety rules.

This example uses level 7
↗
Level 1 + Level 2 + Level 4

Plain-language maintenance log

Turns a plain-language description of work done into a structured log entry, saved with a tool call and linked to the equipment it concerns through a small knowledge graph.

This example uses level 4
↗
Level 0

Keep the household paperwork straight

Organize renewal dates, file names, category totals, and reminders with ordinary code. No model is needed; extracting information from scanned bills is a separate task.

This example uses level 0
↗
Level 1

Turn a meeting transcript into decisions and owners

Turn a transcript into decisions, owners, and open questions in one model call. Someone who attended reviews the draft before it is shared.

This example uses level 1
↗
Level 0 + Level 1 + Level 3

Match invoices to purchase orders

Extract invoice fields, then use code to match purchase orders and compare amounts. Differences go to a person; the model never decides whether the totals reconcile.

This example uses level 3
↗
Level 1 + Level 3

Check an agreement against your own checklist

Check an agreement against a fixed checklist, with cited clauses for each finding. Merge the findings for a person to review.

This example uses level 3
↗
Level 1 + Level 3

Turn an incident write-up into a runbook

Turn an incident write-up into a timeline and repeatable steps. Check owners and success criteria, then ask the incident lead to approve it.

This example uses level 3
↗
Level 1 + Level 3

Turn a script into a shot list

Split a script into scenes and shots, then check that every line is covered and every shot has a source. A person reviews the plan; drawing frames is a separate task.

This example uses level 3
↗
Level 3 + Level 4 + Level 5

Plan a trip and hold the bookings

Checking what is available, what is open and what connects takes a different number of steps every time, which is what level 5 is for. Read-only lookups run unattended; anything that spends money stops for a person, with the price and the cancellation terms in front of them.

This example uses level 5
↗
Level 1 + Level 3 + Level 6

Grade against a rubric, with a second reader

Two independent reviewers apply the same rubric. Disagreements go to the teacher rather than being averaged away.

This example uses level 6
↗
Level 0 + Level 1

Watch a topic for new work and summarize what turns up

Code detects new records from fixed sources. One model call summarizes each new title and abstract; code attaches the original citation. It does not follow references or choose new searches.

This example uses level 1
↗
Level 0 + Level 1

Assemble a weekly status report from several systems

Code assembles the weekly figures; one model call drafts the report. Checks flag unsupported numbers and missing required facts, then a person reviews and sends it.

This example uses level 1
↗
Level 0 + Level 1 + Level 3

Keep a tracker document current from several sources

Keep a shared tracker current through source comparisons and a review queue. Model proposals and changes to human-written fields need approval; missing evidence is flagged.

This example uses level 3
↗
Use case × pattern

34 recipes across 24 techniques

Both domains in one grid. Every technique some recipe uses gets a column, grouped and colored by level; the 30 pages no recipe uses are left out. A filled cell means that recipe uses that technique. The grid scrolls sideways inside its own box; on a narrow screen it is a list instead.

Recipes down the side, techniques across, grouped by level. A filled cell means that recipe uses that technique.
RecipeLevel 0Level 1Level 2Level 3Level 4Level 5Level 6Level 7Topics
When not to use a modelPrompt engineeringStructured outputImages, audio and videoRetrieval-augmented generation (RAG)Knowledge graphs and GraphRAGMemoryPrompt chainingRoutingParallel callsWrite and checkHuman approvalFunction callingCode executionSingle agentAgentic RAG and deep researchCoding agentsSkillsAgent graphsReview and debateAlways-on assistantsEvalsSafety, privacy and governanceOperations
Answer questions about a set of documentsuses Structured outputuses Retrieval-augmented generation (RAG)uses Evals
Sort an inboxuses Structured outputuses Routinguses Human approval
Write a research brief with citationsuses Write and checkuses Agentic RAG and deep research
Coding assistant on your own repouses Coding agentsuses Skillsuses Safety, privacy and governance
Turn photos and PDFs into recordsuses Structured outputuses Images, audio and videouses Human approval
Voice notes into structured entriesuses Images, audio and videouses Prompt chaininguses Evals
Support deskuses Retrieval-augmented generation (RAG)uses Routinguses Human approvaluses Function calling
Nightly source monitoruses Structured outputuses Routinguses Evalsuses Operations
Data analysis by conversationuses Code executionuses Single agent
Drafting with a revieweruses Prompt chaininguses Write and check
A team of personal assistantsuses Memoryuses Agent graphsuses Always-on assistantsuses Safety, privacy and governance
Plain-language maintenance loguses Structured outputuses Knowledge graphs and GraphRAGuses Function calling
Keep the household paperwork straightuses When not to use a model
Turn a meeting transcript into decisions and ownersuses Prompt engineeringuses Structured output
Match invoices to purchase ordersuses When not to use a modeluses Structured outputuses Human approval
Check an agreement against your own checklistuses Structured outputuses Parallel callsuses Human approval
Turn an incident write-up into a runbookuses Structured outputuses Prompt chaininguses Human approval
Turn a script into a shot listuses Structured outputuses Prompt chaining
Plan a trip and hold the bookingsuses Human approvaluses Function callinguses Single agent
Grade against a rubric, with a second readeruses Structured outputuses Human approvaluses Review and debate
Watch a topic for new work and summarize what turns upuses When not to use a modeluses Prompt engineeringuses Structured outputuses Evalsuses Operations
Assemble a weekly status report from several systemsuses When not to use a modeluses Prompt engineeringuses Evalsuses Operations
Keep a tracker document current from several sourcesuses When not to use a modeluses Structured outputuses Human approvaluses Operations
Check measurements against limits, and chart what driftsuses When not to use a model
Sweep a design over its corners and report the marginsuses When not to use a model
Turn a measurement session into a report somebody can reviewuses When not to use a modeluses Prompt engineering
Answer questions from a datasheet, a test spec and a change noticeuses Structured outputuses Retrieval-augmented generation (RAG)
Pull an instrument's accuracy table out of its manualuses When not to use a modeluses Structured outputuses Human approval
Sort failing units and operator notes into causesuses Structured outputuses Routinguses Human approval
Check a board against the design rules documentuses Structured outputuses Write and check
Turn a requirements list into a test planuses Structured outputuses Prompt chaininguses Human approval
Draft an instrument control script from its programming manualuses Retrieval-augmented generation (RAG)uses Write and checkuses Code execution
Ask questions of a production test loguses Function callinguses Code execution
Work a bring-up problem at the benchuses Human approvaluses Function callinguses Single agentuses Safety, privacy and governance
Recipes using it94182411541413533111111534

The columns say which techniques earn their place across many jobs and which are specialized. A recipe's own page gives the reasoning for each choice, and says why it does not need a higher level.