Search

Search the site

Every technique, recipe, teardown, thread, level, glossary term, name and failure mode on the site: 806 entries in all. Nothing typed here is sent anywhere: the whole index loads once and every search runs in your browser.

Type a word to search every technique, recipe, teardown, thread, level, glossary term, name and failure mode on the site. Not sure where to start? The worksheet finds the right level for a task in a few questions, and the glossary defines every term the site uses.

JavaScript is off, or still loading, so here is everything on the site at once, grouped the way search results would be. Use your browser's own find-in-page (usually Ctrl/Cmd-F) to search this list.

Techniques and topics54
When not to use a modelLevel 00 · Conventional software

How to tell when ordinary code, search or a form is enough.

ChatLevel 01 · Direct prompting

Asking a model a question in a chat app.

Prompt engineeringLevel 01 · Direct prompting

Writing instructions that get consistent results.

Structured outputLevel 01 · Direct prompting

Getting answers in a fixed format such as JSON.

Reasoning at answer timeLevel 01 · Direct prompting

Letting the model think for longer before it answers.

Images, audio and videoLevel 01 · Direct prompting

Giving the model images, audio, video and documents, and getting them back.

Context engineeringLevel 02 · Added context

Deciding what goes into the request, and caching the parts that repeat.

Embeddings and searchLevel 02 · Added context

Finding text by meaning instead of by keyword.

Retrieval-augmented generation (RAG)Level 02 · Added context

Searching your documents and giving the results to the model.

Knowledge graphs and GraphRAGLevel 02 · Added context

Storing facts as entities and relations, for questions that span several documents.

MemoryLevel 02 · Added context

Keeping information from one conversation to the next.

Prompt chainingLevel 03 · Workflows

Splitting a task into steps, each with its own prompt.

RoutingLevel 03 · Workflows

Sorting inputs and sending each one to the right prompt.

Parallel callsLevel 03 · Workflows

Running several prompts at once and combining the results.

Write and checkLevel 03 · Workflows

One prompt writes, another checks, and the loop repeats until the check passes.

Workflow graphsLevel 03 · Workflows

Describing a workflow as steps and the connections between them.

Human approvalLevel 03 · Workflows

Pausing for a person to approve or correct.

Function callingLevel 04 · Tool use

Letting the model call functions that you define.

Code executionLevel 04 · Tool use

Letting the model write code and run it in a sandbox.

Model Context ProtocolLevel 04 · Tool use

A standard way to connect models to tools and data.

Computer and browser useLevel 04 · Tool use

Letting the model operate a screen, a mouse and a keyboard.

Single agentLevel 05 · Agent loops

A model that plans, acts and checks its own work in a loop.

The agent harnessLevel 05 · Agent loops

Everything around the model in an agent: the loop, tools, context handling, permissions, caps and sandbox.

Agentic RAG and deep researchLevel 05 · Agent loops

An agent that runs its own searches until it has an answer.

Coding agentsLevel 05 · Agent loops

Agents that read, write, run and test code.

SkillsLevel 05 · Agent loops

Reusable instructions that an agent loads when it needs them.

Voice agentsLevel 05 · Agent loops

Agents you talk to in real time.

Lead agent and workersLevel 06 · Teams of Agents

A lead agent splits the task and hands parts to other agents.

Agent graphsLevel 06 · Teams of Agents

Describing a team of agents and how work passes between them.

Review and debateLevel 06 · Teams of Agents

Agents that check, or argue with, each other's work.

Long-running tasksLevel 07 · Always-on agents

Tasks that run for hours or days.

Always-on assistantsLevel 07 · Always-on agents

Agents that resume work across sessions, schedules, and events.

Organizations of agentsLevel 07 · Always-on agents

Large groups of agents with roles and shared goals.

Robots and machinesLevel 07 · Always-on agents

Models that control robots and other machines.

EvalsTopics at every level

Measuring whether a change made the results better.

Evaluation frameworksTopics at every level

The tools that run test sets and graders for you, and what to check before trusting their numbers.

Changing the modelTopics at every level

Fine-tuning, distillation, synthetic data and automated prompt tuning.

Fine-tuning and adaptersTopics at every level

Training a model further on your own examples, in full or with small adapters such as LoRA.

DistillationTopics at every level

Training a smaller model to reproduce what a larger one does on your task.

Synthetic dataTopics at every level

Using a model to write training or test examples, and checking them before they are used.

Prompt optimizationTopics at every level

Letting a program search for better prompts against a test set.

Safety, privacy and governanceTopics at every level

Prompt injection, permissions, data handling and audit.

GuardrailsTopics at every level

Checks on what goes into a model and what comes out, and the limits of those checks.

Red teamingTopics at every level

Attacking your own system on purpose, before someone else does, and turning what you find into tests.

OperationsTopics at every level

Cost, speed, monitoring and running models on your own hardware.

ObservabilityTopics at every level

Recording what each run did, so a bad result can be traced to the step that caused it.

AI gatewaysTopics at every level

One entry point in front of several model providers, for keys, routing, limits, fallback and logs.

Cost optimizationTopics at every level

Spending fewer tokens and less time for the same result: caching, batching, smaller models, shorter context.

Running models locallyTopics at every level

Running open-weight models on your own hardware: what fits, quantization, and what you give up.

Working with a modelTopics at every level

How to brief a model, review its work and decide what to hand over.

Briefing: saying what you wantTopics at every level

Saying what you want clearly enough that the model does not have to guess.

Reviewing work you did not doTopics at every level

Checking work you did not do yourself before it goes anywhere.

Deciding what to hand overTopics at every level

Deciding which parts of a task to hand to a model and which to keep.

Calibrating trustTopics at every level

Learning, from results over time, how much to rely on a model without checking.

Recipes34
Answer questions about a set of documentsNeeds level 2

Uses RAG, structured output and an eval set. Level 2 is enough because a single search answers most questions.

Sort an inboxNeeds level 3

Sorts mail into fixed categories and produces structured output. A person approves anything that gets sent. The categories are known in advance, so an agent is not needed.

Write a research brief with citationsNeeds level 5

Uses agentic RAG to find sources and a fixed check on every claim against the section it cites. It needs level 5 for the searching; the checking is level 3.

Coding assistant on your own repoNeeds level 5

A coding agent that reads, edits, runs and tests code in your repository, using skills for repeated tasks and a safety review before anything ships.

Turn photos and PDFs into recordsNeeds level 3

Reads the image or PDF, fills a fixed schema, and saves the record once a person confirms it.

Voice notes into structured entriesNeeds level 3

Transcribes a voice note, splits it into steps, and turns each step into a structured entry that an eval set checks for accuracy.

Support deskNeeds level 4

Routes an incoming ticket, searches the documentation for an answer, calls a tool when an action is needed, and hands off to a person when it is unsure.

Nightly source monitorNeeds level 3

Runs on a timer, diffs a set of public pages in code, and asks a model one question about each change. Level 3: the schedule and the checkpoint are infrastructure, not agency.

Data analysis by conversationNeeds level 5

A single agent writes and runs code against a dataset, one question at a time, to answer questions a fixed query could not anticipate.

Drafting with a reviewerNeeds level 3

One prompt drafts a piece of writing and another checks it against a rubric, repeating until the draft passes.

A team of personal assistantsNeeds level 7

Several always-on agents split personal tasks among themselves, sharing memory and staying inside the same safety rules.

Plain-language maintenance logNeeds level 4

Turns a plain-language description of work done into a structured log entry, saved with a tool call and linked to the equipment it concerns through a small knowledge graph.

Keep the household paperwork straightNeeds level 0

Organize renewal dates, file names, category totals, and reminders with ordinary code. No model is needed; extracting information from scanned bills is a separate task.

Turn a meeting transcript into decisions and ownersNeeds level 1

Turn a transcript into decisions, owners, and open questions in one model call. Someone who attended reviews the draft before it is shared.

Match invoices to purchase ordersNeeds level 3

Extract invoice fields, then use code to match purchase orders and compare amounts. Differences go to a person; the model never decides whether the totals reconcile.

Check an agreement against your own checklistNeeds level 3

Check an agreement against a fixed checklist, with cited clauses for each finding. Merge the findings for a person to review.

Turn an incident write-up into a runbookNeeds level 3

Turn an incident write-up into a timeline and repeatable steps. Check owners and success criteria, then ask the incident lead to approve it.

Turn a script into a shot listNeeds level 3

Split a script into scenes and shots, then check that every line is covered and every shot has a source. A person reviews the plan; drawing frames is a separate task.

Plan a trip and hold the bookingsNeeds level 5

Checking what is available, what is open and what connects takes a different number of steps every time, which is what level 5 is for. Read-only lookups run unattended; anything that spends money stops for a person, with the price and the cancellation terms in front of them.

Grade against a rubric, with a second readerNeeds level 6

Two independent reviewers apply the same rubric. Disagreements go to the teacher rather than being averaged away.

Watch a topic for new work and summarize what turns upNeeds level 1

Code detects new records from fixed sources. One model call summarizes each new title and abstract; code attaches the original citation. It does not follow references or choose new searches.

Assemble a weekly status report from several systemsNeeds level 1

Code assembles the weekly figures; one model call drafts the report. Checks flag unsupported numbers and missing required facts, then a person reviews and sends it.

Keep a tracker document current from several sourcesNeeds level 3

Keep a shared tracker current through source comparisons and a review queue. Model proposals and changes to human-written fields need approval; missing evidence is flagged.

Check measurements against limits, and chart what driftsNeeds level 0

Use code to calculate limits, yield, process capability, and trends across lots and fixtures. The pass/fail decision stays deterministic; no model is involved.

Sweep a design over its corners and report the marginsNeeds level 0

Sweep prototype boards across line, load, and temperature. Code calculates margins, uncertainty, and guardbanded verdicts; no model is needed.

Turn a measurement session into a report somebody can reviewNeeds level 1

Turn computed measurements and notebook notes into a report. Code owns the figures, the model writes the prose, and a person checks the finished draft.

Answer questions from a datasheet, a test spec and a change noticeNeeds level 2

Retrieval over the documents an engineer already has, answered with citations that can be checked. The case that matters is a change notice contradicting the datasheet on one number, where the right answer depends on the board revision. Level 2 is enough because one search finds the passage.

Pull an instrument's accuracy table out of its manualNeeds level 3

Extract specification rows from a manual, validate their structure, and calculate uncertainty in code. A person verifies ranges, intervals, and conditions against the source.

Sort failing units and operator notes into causesNeeds level 3

Failing measurements and free-text operator notes are sorted into the causes the failure analysis guide already lists, then routed. A person confirms before anything is scrapped or reworked. The categories are known in advance, so this is classification into fixed classes and not an agent.

Check a board against the design rules documentNeeds level 3

A bill of materials and a netlist summary are checked rule by rule against the written design rules. One pass drafts findings, a second checks each finding against the rule text it cites and drops the ones that cite nothing. Level 3, because code decides every step and the rules do not change between boards. This is the rule check that happens before a review meeting, not the design review report itself: for the report, and the characterization data behind it, see the two recipes this page links in its first paragraph.

Turn a requirements list into a test planNeeds level 3

A fixed chain: read the requirements, propose a test for each, build the traceability table, then check that every requirement has a test and every test names a requirement. A person approves before any of it is adopted. The order of the steps is known in advance, which is what keeps this at level 3.

Draft an instrument control script from its programming manualNeeds level 4

The model drafts commands from the manual for that instrument; code checks every one against the documented command set, runs the script on the simulated instrument, and feeds the errors back for another pass. A person bench-checks before it drives real hardware, and every set point goes through a code-side envelope.

Ask questions of a production test logNeeds level 4

Starts where the dashboard stopped: limits, yield and Cpk are already charted and did not answer the question. The model writes analysis code that runs in a sandbox over the CSV, and a person reads the code as well as the answer. Includes the trap of a column in millivolts under a header that says volts.

Work a bring-up problem at the benchNeeds level 5

An agent with read-only tools, instrument queries, the test log and the datasheet, works a low output down to a cause and proposes the next measurement. Queries run unattended; anything that sets a voltage, a current limit or an output goes through the envelope and a person. Level 5 because each measurement depends on the last.

Teardowns6
A deep-research mode, decodedTeardown · expires 03/17/2027

Decoded into Agentic RAG and deep research, Parallel calls, Review and debate, Lead agent and workers.

A coding agent, decodedTeardown · expires 03/17/2027

Decoded into The agent harness, Single agent, Coding agents, Skills, Lead agent and workers, Safety, privacy and governance.

An always-on agent teammate, decodedTeardown · expires 03/17/2027

Decoded into Always-on assistants, Long-running tasks, Computer and browser use, Human approval.

A search-grounded answer engine, decodedTeardown · expires 03/18/2027

Decoded into Retrieval-augmented generation (RAG), Context engineering, Routing.

A workplace assistant over your own documents, decodedTeardown · expires 03/18/2027

Decoded into Retrieval-augmented generation (RAG), Embeddings and search, Knowledge graphs and GraphRAG, Safety, privacy and governance.

A browser agent, decodedTeardown · expires 03/18/2027

Decoded into Computer and browser use, Single agent, Human approval, Safety, privacy and governance.

Threads4
Graph engineeringThread · across the levels

Two uses of graphs that are often confused: graphs that connect information, and graphs that connect work.

Who approves whatThread · across the levels

The same question asked at every level: which part of this does a person still decide? The answer moves from reading each result to setting the limits a run happens inside.

Checking the workThread · across the levels

How you tell whether it worked, from a person reading one answer to a scored set and a recorded trace. The check changes shape at every level; the question does not.

What the model seesThread · across the levels

One question followed across the site: what is in front of the model this turn, and who decided to put it there. The answers run from a written instruction to a note a session leaves for the next one.

Levels9
Level 00 · Conventional softwareRules, search, and automation

Use ordinary code, search, forms, or a task-specific statistical model when they solve the problem. No generative model is required; classical machine learning can belong here too.

Level 01 · Direct promptingAsk for a response

Give the model instructions and receive a response. A conversation repeats this interaction, with a person directing each turn. Prompting, structured output, reasoning, and multimodal inputs can all fit this pattern.

Level 02 · Added contextSupply relevant information

Add relevant documents, retrieved passages, or stored information to the current request. This supplies context beyond the model’s training without retraining it. Missing, stale, or misleading material can still produce a poor answer.

Level 03 · WorkflowsSoftware organizes the steps

Connect model calls through predefined steps, branches, checks, and retries. A model can classify an input or evaluate a result to route the workflow; software still defines the available paths.

Level 04 · Tool useThe model requests an action

The model can request a search, calculation, code execution, or an action in another application. Software enforces permissions, performs the action, and returns the result. Tool use alone does not create an ongoing agent loop.

Level 05 · Agent loopsObserve, decide, act, repeat

The model uses the goal and observed results to choose an action, revise its approach, or finish. Software executes tools and enforces permissions, approvals, and stopping limits. A run can stop because it is complete, blocked, or out of budget.

Level 06 · Teams of AgentsAgents coordinate work

Agents divide, coordinate, or review work across separate contexts. A coordinator can combine their findings, and the agents may use the same model or different models. Coordination adds overhead, and separate reviewers can still make correlated mistakes.

Level 07 · Always-on agentsResume across sessions and events

Saved state, schedules, and events let an agent start or resume work without a fresh chat message each time. The model need not run continuously, and a dedicated computer or agent team is optional. Permissions, human approvals, monitoring, and stop controls still apply.

Topics at every levelThese topics apply whichever level you use.

Five topics cut across all the levels. Each has its own set of pages.

Views31
Design decisions

Context, workflows, reasoning, agents, tools, persistence, and evaluation tradeoffs

Answer a warranty question with evidence

Retrieve the relevant policy, answer each part of the question, and distinguish an unknown fact from a retrieval miss.

Turn an invoice into a checked record

Extract a useful JSON record, preserve missing fields, and catch a total that does not reconcile.

Build a weekly update without invented progress

Extract evidence into a checked table, then draft an update from that table in a fixed two-call workflow.

Approve the exact change before it happens

Draft a calendar change, bind review to the exact proposal, and detect stale or repeated approvals.

Investigate an incident with bounded tools

Let a model choose read-only diagnostic tools, then require an evidence-backed handoff within six calls.

Resume a monitor without duplicating alerts

Process a stock event, save a local outbox record, and prove that replaying the same event does not create another alert.

Newcomer usability session

Human observation guide and results template. Sessions pending.

A team of agents that improves your project brief

Lead agent and workers: brief writing, parallel receiving agents, independent reviewers, and coordinated revision. Includes supporting evaluation methods.

Worked examples

98 scripted examples across everyday life, engineering, and business.

From a chatbot to an always-on agentView

The evolution of conversational AI through context, workflows, tools, agent loops, optional teams, and persistent operation. Architectural changes explained with diagrams.

Project brief builderView

Describe your task and desired automation for your own AI.

Tools for your projectView

Agent instructions, workflows, audits, tool specifications, checks and handoffs.

Agent instructions builderView

Describe how an agent should work in your project, then let it verify the details against your actual files.

Workflow designerView

Define the complete experience from trigger to delivered result, including your role and exceptions.

Definition-of-done builderView

Turn “it works” into observable checks for quality, completion, and the effort still required from you.

Existing-workflow auditView

Find where an existing process loses time or quality before deciding what to replace.

Tool specification builderView

Describe a reusable tool an agent can build or adapt, including its contract and how to verify it.

Project handoff builderView

Give the next person or agent a factual starting point, without confusing plans with completed work.

MapView

Every technique as a node, laid out by level, with the taxonomy's requires, upgrades-to, combines-with and alternative-to relations as edges.

TimelineView

When each level reached the public, marked at the launch of the product that brought it there, with the paper that first described it behind each mark.

WorksheetView

Answer a few questions about your own job and explore a candidate level, then check automation and human effort.

RecipesView

Common tasks and the levels they need, with the use case by technique matrix underneath.

TeardownsView

Products you have used, taken apart into the techniques they are built from.

NamesView

Every developer, model, product and tool this site names, with the technique each one demonstrates and the date it was checked.

Failure galleryView

Every named failure mode across every technique, grouped by level, each with how to notice it and how to test for it.

GlossaryView

Every term a newcomer meets on this site, defined from the page that explains it.

MethodView

How a level is defined, how a page is written and reviewed, and what the numbers on this site do and do not mean.

Job shapesView

The kinds of job people bring to a model, sorted by the shape of the work rather than its subject, with what moves each one to a lower or a higher level.

What changedView

A dated record of what changed on this site and why it matters to a reader, newest first, with an Atom feed.

For your agentView

The procedure this site hands a reader's own AI assistant: what to ask, how to choose a design that fits the desired outcome, automation, and user effort, and what not to claim.

Glossary106
agentLevel 05 · Agent loops

A model put in charge of a loop rather than one choice: it acts, reads the result, and decides what to do next and when to stop, instead of following steps your code chose in advance.

agent graphLevel 06 · Teams of Agents

A team of agents described as a graph: nodes are agents rather than fixed steps, and the output of one node (usually called the supervisor) picks which agent runs next from a fixed list of names.

agent loopLevel 04 · Tool use

The repeating cycle behind every agent: the model proposes one action from what it currently sees, your code carries it out, and the result goes back to the model, until the model itself decides to stop.

agentic AILevel 05 · Agent loops

A loose umbrella for systems where the model, not your code, chooses each step and decides when the task is finished. On this site that begins at level 5; below it a person or a program picks the steps and the model only fills them in.

agentic RAGLevel 05 · Agent loops

Retrieval where the model, not your code, decides how many times to search, what to search for next, and when it has read enough to answer, instead of searching once and answering once.

AI gatewayTopics at every level

One entry point every call to a model provider goes through instead of calling each provider directly, so provider keys, routing and fallback, per-caller budgets, caching, logging and policy checks sit in one place rather than in every application that calls a model.

always-on agentLevel 07 · Always-on agents

An agent built around a computer of its own that keeps running between the moments a person talks to it, deciding on each scheduled check whether anything needs doing and acting under a fixed approval policy.

approval policyLevel 07 · Always-on agents

A fixed, code-side table that sorts every action a model proposes into one of a few classes (run it automatically, queue it for a person, or never run it), regardless of what the model asked for.

audit trailTopics at every level

A record of what an agent did and why, kept independent of the agent itself, that stands in for a person who was not there to catch a problem as it happened.

automated prompt optimizationTopics at every level

Searching for a better prompt against a measured score instead of a person hand-editing the wording, tuning instructions, examples or weights the way a training run tunes a model.

autonomyThread

How much of a task a system settles for itself. On this site it is not one quality but a question asked level by level: who decides the next step, and what a person is still holding at that level.

batch APITopics at every level

A way to submit many model requests at once for processing that finishes within a day rather than instantly, usually billed at roughly half the price of an ordinary synchronous call.

best-of-NLevel 01 · Direct prompting

Generating several candidate answers to the same question and keeping the one a checker judges best, rather than trusting whatever the first attempt produces.

briefingTopics at every level

Writing down everything a model needs in order to act without guessing: the goal, context it lacks, constraints, what a finished result looks like, what to do when unsure, and the output format.

calibrating trustTopics at every level

Keeping how much you rely on a model without checking in line with how often it has actually been right on tasks like the one in front of you, tracked per task type, not as one overall impression.

capLevel 05 · Agent loops

A hard limit your code enforces on a loop, such as a maximum number of steps or tokens, that forces a stop the model cannot override or even see.

checkpointLevel 03 · Workflows

A saved snapshot of a workflow's shared state, written after a step, so a crashed or interrupted run can resume from that point instead of starting over from the beginning.

chunkLevel 02 · Added context

A passage a document is cut into before indexing, small enough to embed and retrieve on its own, ideally holding one complete idea rather than splitting a sentence or table row in half.

citationLevel 02 · Added context

A pointer from part of an answer back to the specific source passage it came from, so a reader can check whether the source actually supports what the answer claims.

citation hit rateLevel 02 · Added context

The share of questions where every source a grading rule expects was actually cited in the model's answer, used to score how well a retrieval step is working.

code executionLevel 04 · Tool use

Letting the model write a small program instead of choosing among named tools, then running that program in a sandbox your code controls rather than trusting or interpreting it directly.

coding agentLevel 05 · Agent loops

A single agent whose tools read files, edit them, run commands and run tests, looping on a propose-edit-run-test cycle until its own tests pass or a cap ends the run.

compactionLevel 07 · Always-on agents

Replacing an aging context window with a short written summary once it nears its limit, so a session can carry forward what mattered without keeping the full history.

computer useLevel 04 · Tool use

Letting the model operate a real screen: it looks at a screenshot, picks one action such as a click or a keystroke, your code carries it out, and a new screenshot goes back.

content credentialsLevel 01 · Direct prompting

A record of where a piece of media came from, carried with the file as signed assertions, specified by C2PA rather than by one maker. It says whether that history validates and is free from tampering, not whether the history is good or bad.

context engineeringLevel 02 · Added context

Deciding what goes into a model's request (which instructions, examples, documents and history) and in what order, since the model only knows what it was trained on and what the request contains.

context windowLevel 02 · Added context

The amount of text a model can read in one request; material that does not comfortably fit has to be trimmed, retrieved, or summarized before the model ever sees the question.

cosine similarityLevel 02 · Added context

How closely two vectors point in the same direction, used to measure how related two pieces of text are once both have been turned into embeddings.

credential vaultLevel 07 · Always-on agents

Secure storage that holds a password, key or payment method on an agent's behalf, so the agent can use it without ever seeing the raw value itself.

delegatingTopics at every level

Deciding which parts of a task to hand to a model and which to keep, based on what a wrong answer would cost, how checkable the result is, and how reversible the action is.

distillationTopics at every level

Training a smaller model to imitate a larger model's outputs on a given task, so the smaller one can stand in for the larger one on that same narrow job.

embeddingLevel 02 · Added context

A list of floating-point numbers standing in for a piece of text's meaning, positioned so texts with similar meaning get vectors that point in similar directions.

evalTopics at every level

Running the same fixed set of questions against a system before and after a change, graded the same way both times, so a claim that the change helped can be checked instead of assumed.

evaluation frameworkTopics at every level

A tool that holds a dataset of examples, runs a program or a prompt against every one of them, grades each result and lets two runs be compared, instead of a team building that machinery from scratch. Some also record a trace of what happened inside each run.

few-shot exampleLevel 01 · Direct prompting

One or more worked examples included in a prompt to show the model a format or pattern rather than only describing it in words.

fine-tuningTopics at every level

Training a model further on your own examples so its behavior on that kind of task becomes more consistent, without repeating the same instructions in every request.

function callingLevel 04 · Tool use

Giving the model a fixed list of actions your code defined, each with a name, a description and an argument schema, and letting it choose whether to use one, which one, and what arguments to send.

golden setTopics at every level

A fixed list of questions with a known right answer, or a rubric for judging one, run the same way before and after a change so two runs can be fairly compared.

graderTopics at every level

The thing that scores an answer against a golden set, either by matching it exactly against a pattern or by having another model read it against a rubric, which is itself a judgment call worth checking by hand.

graph engineeringThread

One phrase for two different techniques that happen to share a data structure. A knowledge graph connects information: entities and the relationships between them. A workflow graph or an agent graph connects work: steps, and who or what picks the next one.

groundingLevel 02 · Added context

Tying a claim in an answer to a specific source that can actually be checked, the way RAG grounds an answer in retrieved documents instead of whatever the model remembers from training.

guardrailTopics at every level

A check on what goes into a model or what comes out: an input filter, an output validator, a separate classifier trained to judge safety, a schema check. It is probabilistic and can be wrong in both directions, so it is never the control that holds; a code check that tests a specific fact is.

hallucinationTopics at every level

A fluent, confident answer that is not actually true or not supported by any real source; research argues this happens because training and grading reward a plausible guess over admitting uncertainty.

handoffLevel 06 · Teams of Agents

The edge in an agent graph where one agent's output hands control to another named agent, chosen by a model call rather than a rule your code wrote in advance.

harnessLevel 05 · Agent loops

Everything around the model in an agent: the loop that calls it, the tool definitions it is shown and the code that runs them, what goes into its next request, whether an action needs approval, the caps on steps and tokens, the sandbox, and what gets logged. None of it is the model.

human approvalLevel 03 · Workflows

A pause your code inserts before something costly, irreversible or too uncertain to ship, handing the decision to a person instead of letting the run continue on its own.

hybrid searchLevel 02 · Added context

Running a keyword search and a meaning-based search over the same documents and merging the two result lists, so an exact identifier and a paraphrased question can both be found.

indexLevel 02 · Added context

The set of embeddings, or the keyword structure, built from a document set in advance so a later question can be compared against it and ranked, without re-reading every document.

knowledge cutoffLevel 01 · Direct prompting

The date after which a model's training data stops, so anything that changed after that date is not something the model actually knows, however confidently it answers.

knowledge graphLevel 02 · Added context

Facts stored as entities and the relationships between them, so a chain of hops can join facts across documents instead of needing one passage to state the whole answer.

latencyTopics at every level

How long a request takes to complete, tracked separately from cost; caching, batching and routing to a smaller model are among the main ways to bring it down at volume.

least privilegeTopics at every level

Giving a system only the access its task actually needs, so an action it was never granted is one no instruction, however cleverly worded, can talk it into taking.

levelLevel 05 · Agent loops

One of the eight levels this site sorts a technique into. A new level starts where the answer to “who decides the next step” changes: nobody, you, your code, the model for one action, the model for every step, several models, the models including when to start.

local modelTopics at every level

A model run on your own hardware instead of called over an API, trading a per-call bill for hardware you own and buy, and for models small enough to fit on it.

long-running taskLevel 07 · Always-on agents

Work that starts on a schedule or an event and continues across many separate sessions until its queue or goal is finished or a person steps in, with no single session seeing the one before it directly.

loop engineeringLevel 05 · Agent loops

Designing an agent's loop on purpose (what starts it, what it repeats, what stops it) rather than only writing its prompt. Anthropic's Claude Code team defines loops as agents repeating cycles of work until a stop condition is met.

LoRATopics at every level

A small, trainable add-on layered onto a model's frozen weights instead of retraining the whole model, cutting the parameters and memory a fine-tuning run needs by orders of magnitude.

MCPLevel 04 · Tool use

A standard way for an application to connect to servers that expose tools, resources and prompts to a model, instead of a developer wiring each integration by hand.

memoryLevel 02 · Added context

Keeping information from one conversation to the next by writing facts down somewhere and reading them back into a later, otherwise unrelated conversation, since a chat has no memory of its own.

memory layerLevel 02 · Added context

The part of a product that writes facts down after one conversation and reads them back into a later one. Two different things get called this: a summary, cheap to reread but lossy, and a record kept on its own and searched on demand.

model-decided stepLevel 05 · Agent loops

A step in a recorded run where the model's own output chose what happened next (which tool, which query, whether to stop), as opposed to a step your code decided regardless of what the model said.

multi-agent systemLevel 06 · Teams of Agents

Several agents working on a task, using the same underlying model or different models: a lead agent splitting the work among others, agents handing work to each other across a graph, or two agents checking each other's output. That is level 6 on this site.

multimodalLevel 01 · Direct prompting

A model that takes or produces more than text (images, audio, video and documents) inside the same one-call request and response shape as an ordinary chat message.

observabilityTopics at every level

Recording what each run did in enough detail that a bad result can be traced back to the step that caused it: which passages a retrieval step picked, which tool the model called and with what arguments, which branch a workflow took, and what each step spent.

orchestratorLevel 06 · Teams of Agents

A lead model that reads a task, decides how to split it, hands each piece to a worker, and combines what comes back, rather than doing the whole task itself.

organization of agentsLevel 07 · Always-on agents

Several standing agents with distinct roles and a shared goal, where the roster itself keeps running and changing over time rather than being assembled fresh for one job and torn down after.

parallel callsLevel 03 · Workflows

Running more than one model call at the same time instead of one after another, then combining the results in code, either splitting one task into independent parts or running the same task several times to vote.

progressive disclosureLevel 05 · Agent loops

Loading only a skill's short name and description into context up front, and its full instructions only once the model actually decides to use it, so unused skills cost almost nothing.

promptLevel 01 · Direct prompting

The request text sent to a model in one call: instructions, examples and the question, everything the model sees that is not already baked into its training.

prompt cachingLevel 02 · Added context

Reusing a matching, unchanged prefix of a request across calls at a reduced billing rate, which is why makers recommend putting content that never changes first and content that changes every call last.

prompt chainingLevel 03 · Workflows

Splitting one task into a fixed sequence of steps and handing each step's output to the next, with your code deciding how many steps there are and what gate sits between them.

prompt engineeringLevel 01 · Direct prompting

Writing the request itself well (instructions, examples, an assigned role, a required format, sometimes asking the model to reason first) to get a more consistent result from a single call.

prompt injectionTopics at every level

Text written to look like an instruction, hidden in the user's message or in retrieved or tool-returned content, that tries to redirect what the model does instead of answering the actual question.

quantizationTopics at every level

Storing a model's weights at lower precision so a large model fits smaller hardware and may run faster. llama.cpp's own documentation says it shrinks the model and can speed up inference, and that it may cost some accuracy. Name the exact level you ran, not just "4-bit".

queueLevel 07 · Always-on agents

A persisted list of pending work items a long-running task drains one at a time across separate sessions, surviving a crash because the queue is checkpointed after every item, not held only in memory.

RAGLevel 02 · Added context

Searching your own documents for the passages closest to a question, putting those passages in the prompt, and asking the model to answer once using only what it was given.

RAG 2.0Level 02 · Added context

A label rather than a defined technique: different people use it for different bundles of retrieval improvements: better chunking, reranking, hybrid search, retrieval the model itself drives. This site has no page under that name; the pages on RAG, embeddings and search, and agentic RAG cover the substance.

reasoning effortLevel 01 · Direct prompting

How much a model is allowed to work through a problem in its own words before answering; makers recommend more of it for math, debugging and planning and less for simple lookups.

reasoning modelLevel 01 · Direct prompting

A model trained, usually with reinforcement learning, to write out a long chain of thought before it answers, and to notice and fix its own mistakes along the way. OpenAI's o1 (September 2024) was the first sold as one.

recallLevel 02 · Added context

The operation that ranks stored facts or documents against a new question and returns the closest matches, the same mechanism whether the store is a document corpus or a small set of personal facts.

red teamingTopics at every level

Attacking your own system on purpose, before someone else does it without permission, across the system around the model and not only the model itself. What matters is not the report but the change it forces, most durably a test that keeps failing until the finding is fixed.

rerankingLevel 02 · Added context

A second pass over a first, larger cut of retrieved candidates, scored by a model trained specifically to sort results by relevance to the original query.

retrievalLevel 02 · Added context

Searching an index of document chunks for the ones closest to a question, and keeping a fixed number of the best matches to hand to the model.

reviewingTopics at every level

Checking work you did not do yourself before it goes anywhere: pulling out specific claims and checking them against their source, then checking what is missing, then judging the result.

review and debateLevel 06 · Teams of Agents

A separate agent, with its own context and often its own retrieval, checks or argues with another agent's work and decides whether to accept it, instead of one fixed, code-owned test.

routingLevel 03 · Workflows

Looking at an input, deciding which of several fixed kinds it is, and sending it down the handler built for that kind, with a real fallback for whatever fits none of them.

rubricTopics at every level

A written checklist a grader model reads an answer against when the answer is too open-ended to match against a fixed pattern, used to turn a judgment call into a repeatable score.

sandboxLevel 04 · Tool use

An isolated environment with no network access and fixed resource limits, where code the model wrote is actually run, so what the model produces is data your code hands to an interpreter, never code it trusts directly.

schemaLevel 01 · Direct prompting

A fixed shape for a reply (named fields with defined types) that a model's output is constrained to match, so downstream code can parse it without guessing at its structure.

self-consistencyLevel 01 · Direct prompting

Asking the same question several times as independent calls and returning whichever answer the largest share of the samples agree on, instead of trusting a single attempt.

skillLevel 05 · Agent loops

Instructions an agent keeps on the shelf until it decides it needs them, loaded into context only when triggered, instead of text repeated into every single turn like a system prompt.

stateless MCPLevel 04 · Tool use

The current MCP specification, revision 2026-07-28, removed the initialize handshake and protocol-level sessions. Every request now carries its own protocol version and capabilities, and a server that needs state across calls returns an explicit, server-minted handle the client passes back as an ordinary tool argument.

stop conditionLevel 05 · Agent loops

The test a model-driven loop uses to decide it is finished and should give a final answer rather than take another action; a cap can force a stop the model never actually reaches.

structured outputLevel 01 · Direct prompting

A reply constrained to come back in a fixed shape, such as a JSON object with named fields, instead of a paragraph your code has to parse by guessing.

synthetic dataTopics at every level

Training examples generated with a model rather than collected from real use, for a later fine-tuning or evaluation run.

System One modelLevel 01 · Direct prompting

TypeSafe AI's name for a model that generates no text and returns a typed decision with a probability in one fast pass, for use inside software. Jev (September 2026) is the first. The name borrows psychology's fast System 1, as against slow, deliberate System 2.

system promptLevel 01 · Direct prompting

The instructions sent to a model at the start of a request, separate from what the user or the retrieved content says, setting how it should behave for that call.

test-time computeLevel 01 · Direct prompting

Computation spent while answering, as opposed to while training. Reasoning models made it a dial: more thinking time, better answers on hard problems, at the cost of seconds and output tokens.

tokenLevel 02 · Added context

The unit a model's input and output are measured and billed in; every cost strip on this site counts tokens in and tokens out for the run it illustrates.

traceLevel 05 · Agent loops

A recorded, step-by-step account of a run (every model call, every tool call, and whether each step was decided by code or by the model) used to explain how an answer was actually produced.

trajectoryTopics at every level

The path a run took to reach its answer: which tools were called, in what order, with what arguments. Scoring it is a separate measurement from scoring the final answer, and it gives partial credit for the steps a run got right.

vectorLevel 02 · Added context

A fixed-length list of numbers representing a piece of text, an embedding, positioned in space so that texts with related meaning end up close together.

vector databaseLevel 02 · Added context

A database built to hold embeddings and rank them by similarity to a query vector: pgvector, Pinecone, Weaviate and Qdrant are four. It stores the index a retrieval step searches; the embeddings it stores come from an embedding model, not from the database.

vision-language-action modelLevel 07 · Always-on agents

A model that converts vision and language input directly into motor control, letting a robot take a physical action rather than only produce text.

voice agentLevel 05 · Agent loops

A single agent you talk to instead of type to, in real time, that additionally has to decide when a caller has stopped talking and what to do if they interrupt.

write and checkLevel 03 · Workflows

A loop where one prompt writes a draft and a separate prompt checks it against one specific, testable criterion, revising and rechecking until it passes or a fixed revision cap is reached.

Names223
GPT-6 AstraOpenAI · frontier modelAll names →
GPT-5.6 SolOpenAI · frontier modelAll names →
GPT-5.6 TerraOpenAI · mid-size modelAll names →
GPT-5.6 LunaOpenAI · small modelAll names →
gpt-ossOpenAI · open-weight modelAll names →
Claude Opus 5Anthropic · frontier modelAll names →
Claude Sonnet 5Anthropic · mid-size modelAll names →
Claude Haiku 4.5Anthropic · small modelAll names →
Gemini 3.1 ProGoogle · frontier modelAll names →
Gemini 3.8 FlashGoogle · mid-size modelAll names →
Gemma 4Google · open-weight modelAll names →
Grok 4.6SpaceXAI · frontier modelAll names →
Muse Spark 1.3Meta · frontier modelAll names →
Llama 4Meta · open-weight modelAll names →
Qwen3.8Alibaba · open-weight modelAll names →
DeepSeek V4DeepSeek · open-weight modelAll names →
GLM-5.3Z.ai · open-weight modelAll names →
Kimi K3Moonshot AI · open-weight modelAll names →
Mistral Small 4Mistral AI · open-weight modelAll names →
Mistral Large 3Mistral AI · open-weight modelAll names →
Phi-4-miniMicrosoft · small open-weight modelAll names →
Nemotron 3NVIDIA · open-weight modelAll names →
Command A+Cohere · enterprise modelAll names →
WhisperOpenAI · speech-to-text modelAll names →
Sora 2OpenAI · video modelAll names →
Veo 3.1Google · video modelAll names →
Imagen 4Google · image modelAll names →
FLUX 3Black Forest Labs · image and video generation modelAll names →
text-embedding-3OpenAI · embedding modelAll names →
Voyage Embed 4Voyage AI · embedding modelAll names →
Cohere Embed v4Cohere · embedding modelAll names →
Nomic Embed Text v2Nomic · open embedding modelAll names →
BGEBAAI · open embedding modelAll names →
Gemini Robotics 2Google · robotics modelAll names →
GR00TNVIDIA · robotics modelAll names →
π0Physical Intelligence · robotics modelAll names →
DeepSeek-V4.1-FlashDeepSeek · open-weight modelAll names →
π0.7Physical Intelligence · robotics modelAll names →
Stable Diffusion 3.5Stability AI · text-to-image modelAll names →
SunoSuno · text-to-music modelAll names →
Gen-4.5Runway · text-to-video modelAll names →
JevTypeSafe AI · system one decision modelAll names →
ChatGPTOpenAI · chat appAll names →
ClaudeAnthropic · chat appAll names →
GeminiGoogle · chat appAll names →
GrokSpaceXAI · chat appAll names →
Meta AIMeta · chat appAll names →
Microsoft CopilotMicrosoft · chat appAll names →
Mistral VibeMistral AI · ai agent for work and codingAll names →
DeepSeekDeepSeek · chat appAll names →
MidjourneyMidjourney · image generationAll names →
ElevenLabsElevenLabs · voice generationAll names →
Gemini NotebookGoogle · research notebookAll names →
ChatGPT ProjectsOpenAI · files and instructions in a chat appAll names →
Claude ProjectsAnthropic · files and instructions in a chat appAll names →
ChatGPT memoryOpenAI · memory in a chat appAll names →
Claude memoryAnthropic · memory in a chat appAll names →
PerplexityPerplexity · answer engineAll names →
GleanGlean · workplace searchAll names →
Glean Enterprise GraphGlean · knowledge graph inside a workplace search productAll names →
Notion AINotion · workspace assistantAll names →
Microsoft 365 CopilotMicrosoft · workplace assistantAll names →
Microsoft 365 Copilot semantic indexMicrosoft · vector index behind a workplace assistantAll names →
ZapierZapier · automation serviceAll names →
MakeCelonis · automation serviceAll names →
n8nn8n · automation service, self-hostableAll names →
Power AutomateMicrosoft · automation serviceAll names →
DifyLangGenius · visual workflow builderAll names →
FlowiseFlowise · visual workflow builderAll names →
LangflowIBM · visual workflow builderAll names →
GumloopGumloop · visual workflow builderAll names →
ChatGPT data analysisOpenAI · code execution in a chat appAll names →
Custom GPTs with actionsOpenAI · tool calling in a chat appAll names →
Claude connectorsAnthropic · tool connections in a chat appAll names →
Gemini connected appsGoogle · tool connections in a chat appAll names →
Claude CodeAnthropic · coding agentAll names →
CodexOpenAI · coding agentAll names →
CursorAnysphere · coding agent in an editorAll names →
GitHub CopilotGitHub · coding agent in an editorAll names →
JulesGoogle · coding agentAll names →
Antigravity CLIGoogle · coding agentAll names →
Grok BuildSpaceXAI · coding agentAll names →
DevinCognition · coding agentAll names →
Devin DesktopCognition · coding agent in an editorAll names →
Replit AgentReplit · coding agentAll names →
ClineCline · open-source coding agentAll names →
AiderAider · open-source coding agentAll names →
ChatGPT deep researchOpenAI · research agentAll names →
Gemini Deep ResearchGoogle · research agentAll names →
Claude ResearchAnthropic · research agentAll names →
Perplexity Deep ResearchPerplexity · research agentAll names →
Grok DeepSearchSpaceXAI · research agentAll names →
ChatGPT voiceOpenAI · voice assistantAll names →
Gemini LiveGoogle · voice assistantAll names →
Claude Code subagentsAnthropic · multi-agent feature of a coding agentAll names →
Grok HeavySpaceXAI · several agents answering one questionAll names →
Grok BotSpaceXAI · always-on agentAll names →
MuseMeta · always-on agentAll names →
Claude CoworkAnthropic · always-on agentAll names →
ChatGPT WorkOpenAI · always-on agentAll names →
Gemini SparkGoogle · always-on agentAll names →
OpenClawOpenClaw Foundation · always-on agent, self-hostedAll names →
Hermes AgentNous Research · always-on agent, self-hostedAll names →
LinkStripe · digital wallet for checkoutAll names →
ManusManus · general-purpose agentAll names →
OptimusTesla · humanoid robotAll names →
FigureFigure AI · humanoid robotAll names →
NEO1X · home robotAll names →
ZedZed Industries · ai code editorAll names →
Regular expressionsevery language · pattern matchingAll names →
SQLevery database · structured queriesAll names →
ElasticsearchElastic · keyword searchAll names →
OpenSearchopen source · keyword searchAll names →
scikit-learnopen source · classical machine learningAll names →
XGBoostopen source · classical machine learningAll names →
spaCyExplosion · rule-based and statistical text processingAll names →
OpenAI APIOpenAI · model APIAll names →
Claude APIAnthropic · model APIAll names →
Gemini APIGoogle · model APIAll names →
OpenRouterOpenRouter · one API for many modelsAll names →
LiteLLMBerriAI · one API for many modelsAll names →
Cloudflare AI GatewayCloudflare · hosted gateway for model callsAll names →
OllamaOllama · runs models locallyAll names →
llama.cppopen source · runs models locallyAll names →
LM StudioElement Labs · runs models locallyAll names →
vLLMopen source · model serverAll names →
TransformersHugging Face · model libraryAll names →
Instructoropen source · structured output libraryAll names →
Outlinesdottxt · structured output libraryAll names →
PydanticPydantic · schema validationAll names →
LlamaIndexLlamaIndex · retrieval frameworkAll names →
LangChainLangChain · application frameworkAll names →
Haystackdeepset · retrieval frameworkAll names →
pgvectoropen source · vector search in PostgresAll names →
PineconePinecone · vector databaseAll names →
WeaviateWeaviate · vector databaseAll names →
QdrantQdrant · vector databaseAll names →
ChromaChroma · vector databaseAll names →
MilvusZilliz · vector databaseAll names →
FAISSMeta · vector search libraryAll names →
Cohere RerankCohere · rerankerAll names →
GraphRAGMicrosoft · knowledge-graph retrievalAll names →
LazyGraphRAGMicrosoft · knowledge-graph retrievalAll names →
Neo4jNeo4j · graph databaseAll names →
LightRAGopen source · knowledge-graph retrievalAll names →
Mem0Mem0 · memory layerAll names →
ZepZep · memory layerAll names →
LettaLetta · agents with long-term memoryAll names →
LangGraphLangChain · graph frameworkAll names →
Semantic RouterAurelio Labs · library that routes a request by meaningAll names →
TemporalTemporal · durable workflow engineAll names →
PrefectPrefect · workflow engineAll names →
Apache Airflowopen source · workflow engineAll names →
InngestInngest · durable workflow engineAll names →
DSPyStanford NLP · prompt programs and optimizersAll names →
Model Context Protocolopen standard · protocol for tools and dataAll names →
E2BE2B · code sandboxAll names →
ModalModal · code sandbox and computeAll names →
BrowserbaseBrowserbase · hosted browsers for agentsAll names →
PlaywrightMicrosoft · browser automationAll names →
ComposioComposio · prebuilt tool connectionsAll names →
Claude computer useAnthropic · computer-use APIAll names →
OpenAI computer useOpenAI · computer-use APIAll names →
Gemini computer useGoogle · computer-use APIAll names →
Claude Agent SDKAnthropic · agent frameworkAll names →
OpenAI Agents SDKOpenAI · agent frameworkAll names →
Agent Development KitGoogle · agent frameworkAll names →
Pydantic AIPydantic · agent frameworkAll names →
smolagentsHugging Face · agent frameworkAll names →
Agent SkillsAnthropic · format for reusable agent instructionsAll names →
AGENTS.mdopen convention · instructions file for coding agentsAll names →
Gemini Live APIGoogle · voice APIAll names →
LiveKit AgentsLiveKit · voice agent frameworkAll names →
PipecatDaily · voice agent frameworkAll names →
VapiVapi · voice agent platformAll names →
CrewAICrewAI · multi-agent frameworkAll names →
AutoGenMicrosoft · multi-agent frameworkAll names →
Microsoft Agent FrameworkMicrosoft · multi-agent frameworkAll names →
Agent2Agent (A2A) Protocolopen standard · protocol between agentsAll names →
MetaGPTopen source · multi-agent frameworkAll names →
LangSmith DeploymentLangChain · hosting for long-running agentsAll names →
promptfoopromptfoo · eval runnerAll names →
BraintrustBraintrust · eval platformAll names →
LangSmithLangChain · eval and tracing platformAll names →
InspectUK AI Security Institute · eval frameworkAll names →
OpenAI EvalsOpenAI · eval frameworkAll names →
DeepEvalConfident AI · eval frameworkAll names →
Ragasopen source · evals for retrievalAll names →
lm-evaluation-harnessEleutherAI · benchmark runnerAll names →
LangfuseClickHouse · tracing and cost trackingAll names →
HeliconeHelicone · tracing and cost trackingAll names →
OpenTelemetryopen standard · tracing standardAll names →
Llama Guard 4Meta · safety classifierAll names →
NeMo GuardrailsNVIDIA · guardrails frameworkAll names →
Guardrails AIGuardrails AI · guardrails frameworkAll names →
AI Guardrails (Lakera Guard)Check Point · prompt-injection filterAll names →
PyRITMicrosoft · red-teaming frameworkAll names →
UnslothUnsloth · fine-tuning libraryAll names →
Axolotlopen source · fine-tuning libraryAll names →
TRLHugging Face · fine-tuning libraryAll names →
PEFTHugging Face · LoRA and other adaptersAll names →
MLXApple · training and inference on Apple hardwareAll names →
OpenAI fine-tuningOpenAI · hosted fine-tuningAll names →
Together AI fine-tuningTogether AI · hosted fine-tuningAll names →
distilabelArgilla · synthetic dataAll names →
Prompt EngineeringOpenAI · prompt engineering guideAll names →
Prompt Engineering OverviewAnthropic · prompt engineering guideAll names →
Prompt Design StrategiesGoogle · prompt engineering guideAll names →
Prompt GeneratorAnthropic · prompt generation toolAll names →
Prompt ImproverAnthropic · prompt optimization toolAll names →
Batch APIOpenAI · asynchronous batch processing apiAll names →
Message Batches APIAnthropic · asynchronous batch processing apiAll names →
Batch APIGoogle · asynchronous batch processing apiAll names →
Deep AgentsLangChain · agent harnessAll names →
Strands AgentsStrands Agents · agent harness SDKAll names →
AI SDKVercel · TypeScript AI and agent SDKAll names →
Browser UseBrowser Use · browser agent frameworkAll names →
PhoenixArize AI · AI observability and evaluationAll names →
garakNVIDIA · LLM vulnerability scannerAll names →
SGLangSGLang community · model serving runtimeAll names →
Failure modes242
A reward the grader can gameTopics at every level
A handoff outside the allowlistLevel 06 · Teams of Agents
The supervisor never convergesLevel 06 · Teams of Agents
The hop cap ships a thin answerLevel 06 · Teams of Agents
The loop stops on a thin answerLevel 05 · Agent loops
The loop never stops on its ownLevel 05 · Agent loops
No memory beyond what is sentLevel 01 · Direct prompting
Knowledge cutoffLevel 01 · Direct prompting
No way to check its own answerLevel 01 · Direct prompting
Looping without progressLevel 05 · Agent loops
The cap ships a still-broken fixLevel 05 · Agent loops
A cache that never hitsLevel 02 · Added context
Stale static contentLevel 02 · Added context
A superficial checkLevel 06 · Teams of Agents
Chunk boundaries split a factLevel 02 · Added context
Stale indexLevel 02 · Added context
A forbidden-zone target survives clampingLevel 07 · Always-on agents
Grader hackingTopics at every level
A reward the grader can gameTopics at every level
Approval fatigueLevel 03 · Workflows
The threshold is tuned wrongLevel 03 · Workflows
A systematic error looks unanimousLevel 01 · Direct prompting
No answer to extractLevel 01 · Direct prompting
A near-tie decided arbitrarilyLevel 01 · Direct prompting
Stale graphLevel 02 · Added context
Notes drift from what they summarizedLevel 07 · Always-on agents
A crash loses or repeats completed workLevel 07 · Always-on agents
A flagged item never gets resolvedLevel 07 · Always-on agents
Two sessions run at the same timeLevel 07 · Always-on agents
Automation complacencyTopics at every level
Duplicated workLevel 06 · Teams of Agents
Runaway spawningLevel 06 · Teams of Agents
The lead drops a worker at combine timeLevel 06 · Teams of Agents
The team budget ships a partial answerLevel 06 · Teams of Agents
Vocabulary mismatchLevel 00 · Conventional software
A weak match is returned with full confidenceLevel 00 · Conventional software
No synthesis across sectionsLevel 00 · Conventional software
The rule or the index goes stale silentlyLevel 00 · Conventional software
A task is claimed by two roles at onceLevel 07 · Always-on agents
A role quietly exceeds its budgetLevel 07 · Always-on agents
Rate limits under real loadLevel 03 · Workflows
A gate that never failsLevel 03 · Workflows
Instructions that quietly conflictLevel 01 · Direct prompting
One example teaches the wrong lessonLevel 01 · Direct prompting
Tuned on too few casesLevel 01 · Direct prompting
The right passage is not retrievedLevel 02 · Added context
The passage is retrieved but ignoredLevel 02 · Added context
Chunk boundaries split a factLevel 02 · Added context
Stale indexLevel 02 · Added context
Conflicting sourcesLevel 02 · Added context
Confident misrouteLevel 03 · Workflows
Category driftLevel 03 · Workflows
The refusal is not loggedTopics at every level
Looping without progressLevel 05 · Agent loops
Drift from the questionLevel 05 · Agent loops
Tool misuseLevel 05 · Agent loops
Overconfidence at the stopLevel 05 · Agent loops
The wrong skill gets pickedLevel 05 · Agent loops
A loaded skill goes unusedLevel 05 · Agent loops
Schema-valid, still wrongLevel 01 · Direct prompting
A model that ignores the schema anywayLevel 01 · Direct prompting
Retrying on the same mistakeLevel 01 · Direct prompting
A schema stricter than the taskLevel 01 · Direct prompting
The agent talks over the callerLevel 05 · Agent loops
A cycle with no exit conditionLevel 03 · Workflows
The checkpoint is incompleteLevel 03 · Workflows
Timeline97
ELIZAJan 1966 · Level 0

Weizenbaum's paper describes ELIZA, a program that scans input for keywords and applies decomposition and reassembly rules attached to them; the paper says "Keywords and their associated transformation rules constitute the SCRIPT for a particular class of conversation." Code picks the reply, with no model involved.

ElasticsearchFeb 2010 · Level 0

Elastic's own history post places Elasticsearch's first release, which it says "happened to be 0.4.0", in February 2010: keyword search an ordinary developer could run without writing a search engine.

scikit-learnFeb 1, 2010 · Level 0

scikit-learn's own About page says these four "took leadership of the project and made the first public release, February the 1st 2010", packaging classical machine learning algorithms a developer could call without implementing them.

XGBoost: A Scalable Tree Boosting SystemMar 9, 2016 · Level 0

The paper describes a scalable tree boosting system "used widely by data scientists to achieve state-of-the-art results". Classical machine learning, no language model anywhere in it, still advancing in the middle of the deep-learning decade.

Attention Is All You NeedJun 12, 2017 · Level 1

Google researchers introduced the Transformer, an architecture built on attention alone, dispensing with the recurrence and convolutions earlier sequence models relied on. It is the architecture under the one-call language models the rest of this level names.

TransformersOct 29, 2018 · Level 1

GitHub's record of huggingface/transformers gives a creation date of October 29, 2018. The repository describes itself as "the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training": the open library a developer could load a pretrained language model with, years before anyone could buy a chat product.

Better Language Models and Their Implications (GPT-2)Feb 14, 2019 · Level 1

OpenAI's post on GPT-2 reports the model doing reading comprehension, translation and summarization "without any fine-tuning of our models, simply by prompting the trained model in the right way". That is level 1 stated plainly: write the request, read the response, no training step in between.

AI Dungeon 2Dec 5, 2019 · Level 1

A text adventure built on GPT-2: the player types any action in plain English and the model writes what happens next. Its maker's own post is dated Thursday, December 5, 2019. One request, one response, open to anyone, three years before ChatGPT. At first it ran from a shared Google notebook, not an app.

GPT-3: Language Models are Few-Shot LearnersMay 28, 2020 · Level 1

OpenAI's paper describes "an autoregressive language model with 175 billion parameters" applied "without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction". That is the one-call pattern that prompt engineering works within.

The OpenAI APIJun 11, 2020 · Level 1

OpenAI's announcement of June 11, 2020 says “We’re releasing an API for accessing new AI models developed by OpenAI”, providing “a general-purpose ‘text in, text out’ interface”, and that “Today the API runs models with weights from the GPT-3 family”. Access was not open: “we are launching today in a private beta rather than general availability”, with a waitlist.

Chain-of-Thought PromptingJan 28, 2022 · Level 1

The paper reports that "generating a chain of thought -- a series of intermediate reasoning steps -- significantly improves the ability of large language models to perform complex reasoning", and that it "improves performance on a range of arithmetic, commonsense, and symbolic reasoning tasks". Nothing about the model changes; only the prompt does.

InstructGPTMar 4, 2022 · Level 1

OpenAI fine-tuned GPT-3 on human demonstrations and human rankings of its outputs. The paper reports that "outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters", in human evaluations on OpenAI's own prompt distribution.

ChatGPTNov 30, 2022 · Level 1

OpenAI's announcement, dated November 30, 2022, says “We’ve trained a model called ChatGPT which interacts in a conversational way” and “During the research preview, usage of ChatGPT is free. Try it now at chat.openai.com.” A chat product anyone could open in a browser, with no API key, no code and no invitation.

GPT-4 Technical ReportMar 15, 2023 · Level 1

OpenAI's technical report describes "a large-scale, multimodal model which can accept image and text inputs and produce text outputs", and reports that it "exhibits human-level performance on various professional and academic benchmarks". That is OpenAI's own evaluation of its own model.

GPT-4oMay 13, 2024 · Level 1

OpenAI's API changelog entry for May 13, 2024 reads: "Released GPT-4o in the API. GPT-4o is our fastest and most affordable flagship model." The changelog does not describe how its modalities are combined; this site cites it only for the release date and OpenAI's own description.

Structured OutputsAug 6, 2024 · Level 1

OpenAI's API changelog entry for August 6, 2024 reads: "Launched Structured Outputs—model outputs now reliably adhere to developer supplied JSON Schemas."

OpenAI o1, the first reasoning modelSep 12, 2024 · Level 1

OpenAI's post of September 12, 2024 opens: "We are introducing OpenAI o1, a new large language model trained with reinforcement learning to perform complex reasoning. o1 thinks before it answers—it can produce a long internal chain of thought before responding to the user." It adds that the model "learns to recognize and correct its mistakes" and that its performance improves "with more time spent thinking (test-time compute)". The extra work happens inside one request and response, which is why this sits at level 1 and not higher. OpenAI's API changelog for the same day records the release of o1-preview and o1-mini.

DeepSeek-R1Jan 22, 2025 · Level 1

The paper reports that "the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories", and that the training brings out "advanced reasoning patterns, such as self-reflection, verification, and dynamic strategy adaptation." A second lab, in the open, reaching what o1 had shown four months earlier.

Gemini 3.1 ProFeb 19, 2026 · Level 1

Google's Gemini API changelog entry for February 19, 2026 reads: "Released Gemini 3.1 Pro Preview, our latest iteration in the new Gemini 3 series family."

GPT-6 AstraSep 3, 2026 · Level 1

OpenAI's API changelog entry for September 3, 2026 reads: "Released GPT-6 Astra, our most capable model, built for the hardest end-to-end work." That is OpenAI's own description of its own model.

DeepSeek-V4.1-FlashSep 10, 2026 · Level 1

DeepSeek's release note of September 10, 2026 says "V4.1-Flash is now live on the DeepSeek API with native multimodal support", describing a mixture-of-experts model with 552B parameters and 8B active for input, 16B for output: a frontier-class open-weight release, one call at a time.

Jev, a typed decision model (early access)Sep 15, 2026 · Level 1

TypeSafe's founder, Diogo Almeida, announces "a new class of frontier models built to make fast, structured decisions that software can use directly": "unstructured state in, typed probabilistic decisions out." Jev writes no text. Its possible outputs are defined in advance, every answer carries a probability, and all of it comes back in one parallel pass, not word by word. TypeSafe quotes "70ms-500ms" a call and "$0.042 / MTok" of input. Its own list of uses is this site's level 3: "classify, route, score, extract, or branch where hand-written logic is too brittle."

FAISSFeb 7, 2017 · Level 2

GitHub's record of facebookresearch/faiss gives a creation date of February 7, 2017 and describes it as "A library for efficient similarity search and clustering of dense vectors". That is the open library a developer could build vector search over their own documents with, and it is still the substrate under most retrieval code.

Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksMay 22, 2020 · Level 2

The paper introduces RAG, in which "the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index of Wikipedia", "accessed with a pre-trained neural retriever": the model answers from passages retrieved for the question rather than only from its own weights.

Embeddings in the OpenAI APIJan 25, 2022 · Level 2

OpenAI's post of January 25, 2022 says “We are introducing embeddings, a new endpoint in the OpenAI API that makes it easy to perform natural language and code tasks like semantic search, clustering, topic modeling, and classification”, and that “Embeddings that are numerically similar are also semantically similar.” Vectors as a service: a developer no longer had to train or host an embedding model.

LlamaIndexNov 2, 2022 · Level 2

GitHub's record of the repository now published as run-llama/llama_index gives a creation date of November 2, 2022, and the first release of its package on PyPI, then named gpt-index, is dated November 22, 2022 with the summary “Building an index of GPT summaries.” The first widely used open library built for this level's job: index a set of documents, retrieve from it, and hand what comes back to a language model.

NeevaAIJan 6, 2023 · Level 2

Neeva's post, bylined "The Neeva Team on 01/06/23", introduces a written answer at the top of its search results with citations embedded in the text, open at once to account holders in the United States. A month before the new Bing. Neeva closed its search engine later in 2023.

Bing Chat launches (the new Bing)Feb 7, 2023 · Level 2

Microsoft's announcement says "Bing reviews results from across the web to find and summarize the answer you're looking for" and that "The new Bing also cites all its sources". On this date it was a limited preview on desktop with a waitlist, so it is not the level's availability date.

the new Bing, open to everyoneMay 4, 2023 · Level 2

Microsoft's announcement says "the new Bing is now in Open Preview and no longer has a waitlist": anyone with a Microsoft account could now ask a question and get an answer written from pages retrieved for it, with links to those pages.

100K context windowsMay 11, 2023 · Level 2

Anthropic's announcement says: "We've expanded Claude's context window from 9K to 100K tokens, corresponding to around 75,000 words!" Enough room to paste the material in rather than retrieve from it.

text-embedding-3Jan 25, 2024 · Level 2

OpenAI's API changelog entry for January 25, 2024 reads: "Released embedding V3 models and an updated GPT-4 Turbo preview". The changelog entry itself says nothing further about the models; this site cites it for the date only.

GraphRAGFeb 13, 2024 · Level 2

Microsoft Research's post describes an approach in which "The LLM processes the entire private dataset, creating references to all entities and relationships within the source data, which are then used to create an LLM-generated knowledge graph", which is then clustered and pre-summarized so questions can be answered across many documents rather than from the nearest few passages.

ChatGPT memoryFeb 13, 2024 · Level 2

OpenAI's post of February 13, 2024 says “We’re testing the ability for ChatGPT to remember things you discuss to make future chats more helpful”, that a user “can explicitly tell it to remember something, ask it what it remembers, and tell it to forget conversationally or through settings”, and that memory can be turned off entirely.

Gemini 1.5 ProFeb 15, 2024 · Level 2

Google's announcement says "We can now run up to 1 million tokens in production", and that at this date "a limited group of developers and enterprise customers can try it with a context window of up to 1 million tokens". A private preview, not a general release, and the standard window was 128,000 tokens.

Prompt cachingOct 1, 2024 · Level 2

OpenAI's API changelog entry for October 1, 2024 reads: "Prompt caching: Discounts and faster processing times on recently seen input tokens." Anthropic's own API release notes record prompt caching leaving beta on the Claude API on December 17, 2024. Reusing a long shared prefix is what makes a large fixed context affordable to send on every request.

NotebookLM becomes Gemini NotebookJul 16, 2026 · Level 2

Google's announcement says: "We're renaming NotebookLM to Gemini Notebook. It's the same standalone product, now doing more across the Google ecosystem and updated with a secure cloud computer." The product whose whole premise is answering only from the documents you gave it is still the plainest consumer example of level 2.

AI ChainsOct 4, 2021 · Level 3

The paper proposes "Chaining LLM steps together, where the output of one step becomes the input for the next, thus aggregating the gains per step", and is titled for what that buys: transparent and controllable human-AI interaction, because the intermediate results exist as steps a person can see.

LangChainOct 25, 2022 · Level 3

PyPI's release history for the langchain package shows version 0.0.1 uploaded on October 25, 2022, with 0.0.2 the next day: the library a developer could install instead of writing chaining and model-swapping plumbing themselves.

OpenAI steps in ZapierDec 9, 2022 · Level 3

Zapier's own OpenAI integration page offers an action that “Sends a prompt to OpenAI and generate a response” inside a multi-step Zap, and says “Zapier lets you connect OpenAI with thousands of the most popular apps, so you can automate your work and have more time for what matters most—no code required.” The flow decides what runs next; one of its steps calls a model. The page carries no launch date, so the date marked here is the earliest capture of it this site could read: the feature was live by then, and may well have shipped earlier.

Microsoft 365 Copilot unveiledMar 16, 2023 · Level 3

Microsoft's announcement says Copilot "is more than OpenAI’s ChatGPT embedded into Microsoft 365. It’s a sophisticated processing and orchestration engine working behind the scenes to combine the power of LLMs, including GPT-4, with the Microsoft 365 apps and your business data in the Microsoft Graph". Software runs the steps (fetch the person's files and mail, build the prompt, call the model, check the result, write into Word or Outlook) and the model fills them in. Shown that day: a first draft in Word from your own files, a deck in PowerPoint from a document, trend analysis in Excel, thread summaries and draft replies in Outlook, and live meeting summaries in Teams.

Microsoft 365 Copilot on saleNov 1, 2023 · Level 3

Microsoft's post of November 1, 2023 opens: "Starting today, Microsoft 365 Copilot is generally available for enterprise customers worldwide."

Semantic RouterNov 9, 2023 · Level 3

PyPI shows version 0.0.1 of semantic-router uploaded on November 9, 2023. Its project page calls it "a superfast decision-making layer for your LLMs and agents" that routes requests "using semantic meaning" rather than waiting on a model's generation: code picks the branch.

AI Builder GPT Prompts in Power AutomateDec 19, 2023 · Level 3

Microsoft's Power Platform blog says "GPT Prompts with Prompt Builder, a new feature of AI Builder, is now generally available!" and that it lets a person "add content processing and content generation capabilities to Power Automate". An ordinary customer drops a model call into a flow whose next step is still chosen by the flow, not the model.

LangGraphJan 17, 2024 · Level 3

LangChain's launch post says "LangGraph is module built on top of LangChain to better enable creation of cyclical graphs, often needed for agent runtimes", and that until then "we've lacked a method for easily introducing cycles into these chains." A developer describes the application as a state machine of nodes and edges.

Batch APIApr 15, 2024 · Level 3

OpenAI's API changelog records the Batch API's release on April 15, 2024. Sending a set of independent requests together, and collecting them when they finish, is the parallel-calls pattern offered as a product feature rather than assembled by hand.

interrupt: human approval in LangGraphDec 14, 2024 · Level 3

LangChain's post introduces interrupt, which will "pause execution of the graph, mark the thread you are running as interrupted, and put whatever you passed as an input to interrupt into the persistence layer." The pattern it names first: "Pause the graph before a critical step, such as an API call, to review and approve the action. If the action is rejected, you can prevent the graph from executing the step".

Building Effective AgentsDec 19, 2024 · Level 3

Anthropic's engineering post names and diagrams five workflow patterns (prompt chaining, routing, parallelization, orchestrator-workers and evaluator-optimizer) and advises: "Start with simple prompts, optimize them with comprehensive evaluation, and add multi-step agentic systems only when simpler solutions fall short."

Flowise sunsetAug 10, 2026 · Level 3

Flowise's sunset notice says "we've decided to wind down our operations for Flowise", with a feature freeze on July 29, 2026 and the repository archived on August 10, 2026, adding that "Flowise source code will still remain on Github and the Apache 2.0 licensed code is yours to keep building on." A visual workflow builder closing while the pattern it built stayed in use.

MRKL SystemsMay 1, 2022 · Level 4

The paper describes a language model surrounded by tools (a calculator, a currency converter, a database call) and "a router that routes every incoming natural language input to a module that can best respond to the input". The router is itself a small neural network. A model chooses the tool and code runs it: the whole paper is about that one idea, nine months before Toolformer.

ToolformerFeb 9, 2023 · Level 4

The paper describes a model trained in a self-supervised way, from a handful of examples per API, to "decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction": the model choosing the action, not code choosing it for the model.

ChatGPT pluginsMar 23, 2023 · Level 4

OpenAI's post of March 23, 2023 says “We’ve implemented initial support for plugins in ChatGPT”, tools that “help ChatGPT access up-to-date information, run computations, or use third-party services”, and that “we're also hosting two plugins ourselves, a web browser and code interpreter”. The model picks which plugin to call mid-conversation. This was an invitation: “Today, we will begin extending plugin alpha access to users and developers from our waitlist.”

ChatGPT plugins open to Plus subscribersMay 12, 2023 · Level 4

OpenAI's ChatGPT release notes carry an entry headed “Web browsing and Plugins are now rolling out in beta (May 12)”, which says “If you are a ChatGPT Plus user, enjoy early access to experimental new features” through a beta panel “which is rolling out to all Plus users over the course of the next week”, and describes plugins as “a new version of ChatGPT that knows when and how to use third-party plugins that you enable”. No waitlist and no invitation: a subscription was enough.

Function callingJun 13, 2023 · Level 4

OpenAI's post of June 13, 2023 says “Developers can now describe functions to gpt-4-0613 and gpt-3.5-turbo-0613, and have the model intelligently choose to output a JSON object containing arguments to call those functions”, through “new API parameters in our /v1/chat/completions endpoint, functions and function_call”. The model chooses the call; the developer's code makes it.

Bard ExtensionsSep 19, 2023 · Level 4

Google's announcement says Extensions let Bard "find and show you relevant information from the Google tools you use every day — like Gmail, Docs, Drive, Google Maps, YouTube, and Google Flights and hotels — even when the information you need is across multiple apps and services." The model decides which of those to reach for mid-conversation; Google's code makes the call. The product is now Gemini's connected apps.

Computer useOct 22, 2024 · Level 4

Anthropic's announcement describes Claude working a computer "by looking at a screen, moving a cursor, clicking buttons, and typing text", and says of the capability: "At this stage, it is still experimental—at times cumbersome and error-prone." It adds that "we encourage developers to begin exploration with low-risk tasks."

Claude analysis toolOct 24, 2024 · Level 4

Anthropic's announcement describes a tool that lets Claude "write and run JavaScript code directly in Claude.ai" to "process data, conduct analysis, and produce real-time insights", available to "all Claude.ai users in feature preview": code execution reaching people who write none.

Model Context ProtocolNov 25, 2024 · Level 4

Anthropic's announcement says "The Model Context Protocol is an open standard that enables developers to build secure, two-way connections between their data sources and AI-powered tools", and that it provides "a universal, open standard for connecting AI systems with data sources, replacing fragmented integrations with a single protocol."

Responses APIMar 11, 2025 · Level 4

OpenAI's API changelog for March 11, 2025 records "the Responses API, a new API for creating and using agents and tools" and, the same day, "a set of built-in tools" for it: web search, file search and computer use. Remote MCP servers and a code interpreter were added on May 20, 2025.

ReActOct 6, 2022 · Level 5

The paper interleaves reasoning traces with actions, so that "reasoning traces help the model induce, track, and update action plans as well as handle exceptions, while actions allow it to interface with external sources, such as knowledge bases or environments, to gather additional information." The model reads each result and picks the next move, which is what makes the loop the model's rather than the code's.

AutoGPTMar 16, 2023 · Level 5

GitHub's record of the Significant-Gravitas/AutoGPT repository gives a creation date of March 16, 2023. The project describes itself as "The open-source platform for AI agents" and says it "lets you build, deploy, and run AI agents that carry out complete workflows": a goal-seeking loop a developer could run instead of writing one.

AgentGPTApr 9, 2023 · Level 5

A free web page, live by April 9, 2023: "Assemble, configure, and deploy autonomous AI Agents in your browser. Create an agent by adding a name / goal, and hitting deploy!" It wrote itself a task list, worked through it and added tasks from the results. Part of the AutoGPT wave of spring 2023, which put a looping agent in front of anyone with a browser, and which few people remember as a product.

DevinMar 12, 2024 · Level 5

Cognition's announcement, headed "Introducing Devin, the first AI software engineer", says it equipped Devin "with common developer tools including the shell, code editor, and browser within a sandboxed compute environment—everything a human would need to do their work." On this date it was not something a customer could buy: "Devin is currently in early access as we ramp up capacity."

Replit AgentSep 5, 2024 · Level 5

Replit's post says “Last week, we launched Replit Agent, our AI system that can create and deploy applications”, and that “It configures your development environment, installs dependencies, and executes code”: the model choosing each next step and stopping when the app runs. Replit adds: “The agent is available today in early access to all Replit Core subscribers” and “it should be treated as ‘alpha’ software.” A subscription, not an invitation.

Realtime APIOct 1, 2024 · Level 5

OpenAI's API changelog entry for October 1, 2024 reads: "Realtime API: Build fast speech-to-speech experiences into your applications using a WebSockets interface." The same changelog records it becoming generally available on August 28, 2025.

Devin, generally availableDec 10, 2024 · Level 5

Cognition's post of December 10, 2024, headed “Devin is now generally available”, says “Today we’re making Devin generally available starting at $500 a month for engineering teams”. The nine months between this and Devin's announcement are the gap between a demonstration and something a customer could buy.

Deep Research in GeminiDec 11, 2024 · Level 5

Google's announcement says Deep Research creates "a multi-step research plan for you to either revise or approve", then works "browsing the web the way you do: searching, finding interesting pieces of information and then starting a new search based on what it's learned", ending in "a comprehensive report of the key findings". It launched that day on desktop and mobile web for Gemini Advanced subscribers.

OperatorJan 23, 2025 · Level 5

OpenAI's post introduces "A research preview of an agent that can use its own browser to perform tasks for you." The person describes a task; the model looks at screenshots of a browser running in the cloud and clicks and types until the job is done. It hands back when it should: "Operator is trained to proactively ask the user to take over for tasks that require login, payment details, or when solving CAPTCHAs." The post says "Available to Pro users in the U.S."

ChatGPT deep researchFeb 2, 2025 · Level 5

OpenAI's post of February 2, 2025 describes deep research as “An agent that uses reasoning to synthesize large amounts of online information and complete multi-step research tasks for you”, which “conducts multi-step research on the internet for complex tasks” and returns a report with citations. The same page says: “Available to Pro users today, Plus and Team next.”

Claude CodeFeb 24, 2025 · Level 5

Anthropic's announcement says Claude Code, released "as a limited research preview" alongside Claude 3.7 Sonnet, "enables developers to delegate substantial engineering tasks to Claude directly from their terminal", describing it as "an active collaborator that can search and read code, edit files, write and run tests, commit and push code to GitHub, and use command line tools".

Claude ResearchApr 15, 2025 · Level 5

Anthropic's announcement says Claude "operates agentically, conducting multiple searches that build on each other while determining exactly what to investigate next", across "both your internal work context and the web". It launched "in early beta for Max, Team, and Enterprise plans in the United States, Japan, and Brazil."

ChatGPT agentJul 17, 2025 · Level 5

OpenAI's post of July 17, 2025 says “ChatGPT can now do work for you using its own computer, handling complex tasks from start to finish”, with “a visual browser that interacts with the web through a graphical-user interface, a text-based browser for simpler reasoning-based web queries, a terminal, and direct API access”, and that “Starting today, Pro, Plus, and Team users can activate ChatGPT’s new agentic capabilities”. The person still starts every task, which is what keeps this at level 5 rather than level 7.

Agent SkillsOct 16, 2025 · Level 5

Anthropic's announcement says "Skills are folders that include instructions, scripts, and resources that Claude can load when needed", and that "Claude will only access a skill when it's relevant to the task at hand": the agent deciding which of its own instructions to read.

Gemini CLI becomes Antigravity CLIMay 19, 2026 · Level 5

Google's post says "we're unifying our efforts into Google Antigravity, our premier agent-first development platform", and that "On June 18, 2026, Gemini CLI and Gemini Code Assist IDE extensions will stop serving requests for Google AI Pro and Ultra, as well as those using it free of charge using Gemini Code Assist for individuals." Enterprise licenses keep Gemini CLI.

CAMELMar 31, 2023 · Level 6

Two model agents, one given the role of the person with a task and one the role of the assistant, work the task out between them with no human in the conversation. The abstract proposes "a novel communicative agent framework named role-playing" and reports "comprehensive studies on instruction-following cooperation in multi-agent settings." Eight weeks before the multi-agent debate paper.

Improving Factuality and Reasoning in Language Models through Multiagent DebateMay 23, 2023 · Level 6

The paper has "multiple language model instances propose and debate their individual responses and reasoning processes over multiple rounds to arrive at a common final answer", and reports that this "improves the factual validity of generated content, reducing fallacious answers and hallucinations that contemporary models are prone to".

AutoGenAug 18, 2023 · Level 6

GitHub's record of the microsoft/autogen repository gives a creation date of August 18, 2023; the first pyautogen release reached PyPI a week later, on August 25, 2023. The README calls it "a framework for creating multi-agent AI applications that can act autonomously or work alongside humans".

GensparkJun 18, 2024 · Level 6

A search product that builds a page of results for each question. Its launch post, dated Jun 18, 2024, describes a "multi-agent framework" and a "team of specialized AI agents". It is the earliest product this site found whose own launch page says several agents share one job, ten months before Claude Research. The post never says how the agents divide the work.

Agent2Agent ProtocolApr 9, 2025 · Level 6

Google's announcement says "The A2A protocol will allow AI agents to communicate with each other, securely exchange information, and coordinate actions on top of various enterprise platforms or applications", and says more than 50 technology partners contributed to it.

Claude Research reaches customersApr 15, 2025 · Level 5

Anthropic's launch post of April 15, 2025 says Claude “operates agentically, conducting multiple searches that build on each other while determining exactly what to investigate next”, and that “Research is now available in early beta for Max, Team, and Enterprise plans in the United States, Japan, and Brazil.” Anthropic's engineering post two months later says of the same feature: “Our Research system uses a multi-agent architecture with an orchestrator-worker pattern, where a lead agent coordinates the process while delegating to specialized subagents that operate in parallel.” The launch post itself does not say several agents run on one request.

How we built our multi-agent research systemJun 13, 2025 · Level 6

Anthropic's engineering post says of a feature customers were already using: "Our Research system uses a multi-agent architecture with an orchestrator-worker pattern, where a lead agent coordinates the process while delegating to specialized subagents that operate in parallel." The subagents work "with their own context windows" before condensing what they found for the lead agent.

Grok 4 HeavyJul 9, 2025 · Level 6

xAI's announcement says: "We have made further progress on parallel test-time compute, which allows Grok to consider multiple hypotheses at once. We call this model Grok 4 Heavy", sold through "a new SuperGrok Heavy tier". Several runs on one question, bought as a tier.

Generative AgentsApr 7, 2023 · Level 7

The paper puts "a small town of twenty five agents" in a sandbox. Its architecture, in the paper's words, is one that agents use to "store a complete record of the agent's experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior": agents that decide when to act, not only what to answer.

π0Oct 31, 2024 · Level 7

Physical Intelligence's post says "We've developed a general-purpose robot foundation model that we call π0 (pi-zero)" that "spans images, text, and actions and acquires physical intelligence by training on embodied experience from robots, learning to directly output low-level motor commands", trained on open-source data plus dexterous tasks collected "across 8 distinct robots".

ManusMar 6, 2025 · Level 7

Manus describes itself, on its own site, as building "general AI agents as the Action Engine for life" and as "building the hands for AI to do" rather than the reasoning underneath: an agent a person points at a goal and lets run.

Gemini RoboticsMar 12, 2025 · Level 7

Google DeepMind's announcement introduces "Gemini Robotics, an advanced vision-language-action (VLA) model that was built on Gemini 2.0 with the addition of physical actions as a new output modality for the purpose of directly controlling robots", alongside Gemini Robotics-ER for spatial reasoning. The registry's Gemini Robotics entry is the later Gemini Robotics 2.

Hermes AgentJul 22, 2025 · Level 7

GitHub's record of NousResearch/hermes-agent gives a creation date of July 22, 2025 and an MIT license. The README calls it “The self-improving AI agent built by Nous Research” and says “Run it on a $5 VPS, a GPU cluster, or serverless infrastructure that costs nearly nothing when idle”; the project's own site (hermes-agent.nousresearch.com) says “Tasks that need an active agent will not run while it is stopped; hosted agents are managed separately in Nous Portal.” That is standing work that runs because the agent is running, on hardware the developer keeps up.

OpenClawNov 24, 2025 · Level 7

GitHub's record of openclaw/openclaw gives a creation date of November 24, 2025. Its README says “OpenClaw is an open-source AI assistant that runs on your own computer and meets you in the channels you already use”, with “One Gateway” running it “as a personal assistant on a laptop or as a shared team deployment”, and carries an MIT license badge.

Claude CoworkJan 12, 2026 · Level 5

Anthropic's dated release notes record, on January 12, 2026, “Cowork research preview on Claude Desktop (macOS only) for Max plans”, and say Cowork “brings Claude Code’s agentic capabilities to the Claude desktop app for knowledge work beyond coding”. Pro plans followed on January 16, 2026 and the same notes record “Claude Cowork is now generally available on macOS and Windows through the Claude Desktop app” on April 9, 2026.

π0.7Apr 16, 2026 · Level 7

Physical Intelligence's post calls π0.7 "a steerable generalist model that can perform dexterous tasks across robots, scenes, and skills", and reports "compositional generalization, recombining skills from various tasks to solve new problems", including a robot folding laundry with no laundry-folding data in its training.

Gemini SparkMay 19, 2026 · Level 7

Google's announcement introduces "Gemini Spark, a 24/7 personal AI agent that helps you navigate your digital life", "deeply integrated with the Workspace tools you rely on daily, like Gmail, Docs, Slides and more", and says that "because it is a cloud-based agent, Spark keeps working in the background even when you close your laptop or lock your phone."

Gemini Spark in subscribers' handsJun 30, 2026 · Level 7

Google's post of June 30, 2026 says "Gemini Spark for macOS is available in Beta to Google AI Ultra subscribers aged 18 and over, starting in the US." The May announcement had described Spark as "a 24/7 personal AI agent" that takes "recurring tasks or triggers" and "keeps working in the background even when you close your laptop or lock your phone."

Claude Cowork runs without your computerJul 7, 2026 · Level 7

Anthropic's dated release notes record, on July 7, 2026: "Cowork runs your sessions remotely (in beta), so your sessions and files are saved to your Claude account and go where you go, on any device. Work continues when you close your laptop, and scheduled tasks run with no device online." This is the entry that makes Cowork an always-on product; the January release ran on the person's own machine.

ChatGPT WorkJul 9, 2026 · Level 7

OpenAI's product page says “Powered by GPT‑5.6, ChatGPT Work brings together context from your team’s tools to turn scattered notes, drafts, and ideas into finished work — and keeps projects moving while you stay in control”, and that it “gathers context, plans the approach, and takes action across your tools, files, and desktop apps”.

Grok BotAug 11, 2026 · Level 7

xAI's announcement says "Grok Bot is your team of always-on agents" and, of those agents, "They have their own computer, work inside tools and apps like you do, and keep working 24/7." It says Grok Bot "is in beta and available today" to named SuperGrok and Cursor subscriber tiers on desktop and iOS.

MuseSep 8, 2026 · Level 7

Meta's announcement says "Muse is a personal AI agent. It doesn't just answer questions, it actually does the work", and that "A separate Sentinel agent runs on that same machine, kept apart from Muse at the system level. Nothing Muse does reaches the internet unless the Sentinel approves it, and it asks the person for permission when needed." A checking agent shipped inside a consumer product, not only proposed in a paper.

Cowork merges into ClaudeSep 16, 2026 · Level 7

Anthropic's announcement folds a separate always-on product back into the chat app: "hand over a report due at noon, and Claude takes it from there, even after you've closed your laptop", and "You can check progress from your phone on the way to the office." Cowork's capabilities are now "available from any conversation, with the context, skills, and connectors you already have."

Perplexity AskDec 7, 2022 · Level 2

The early Perplexity Ask generated answers from retrieved search results with citations. Cofounder Aravind Srinivas retrospectively dates its launch to December 7, 2022 at Stripe Sessions 2024.