Search the site
Every technique, recipe, teardown, thread, level, glossary term, name and failure mode on the site: 806 entries in all. Nothing typed here is sent anywhere: the whole index loads once and every search runs in your browser.
JavaScript is off, or still loading, so here is everything on the site at once, grouped the way search results would be. Use your browser's own find-in-page (usually Ctrl/Cmd-F) to search this list.
How to tell when ordinary code, search or a form is enough.
Asking a model a question in a chat app.
Writing instructions that get consistent results.
Getting answers in a fixed format such as JSON.
Letting the model think for longer before it answers.
Giving the model images, audio, video and documents, and getting them back.
Deciding what goes into the request, and caching the parts that repeat.
Finding text by meaning instead of by keyword.
Searching your documents and giving the results to the model.
Storing facts as entities and relations, for questions that span several documents.
Keeping information from one conversation to the next.
Splitting a task into steps, each with its own prompt.
Sorting inputs and sending each one to the right prompt.
Running several prompts at once and combining the results.
One prompt writes, another checks, and the loop repeats until the check passes.
Describing a workflow as steps and the connections between them.
Pausing for a person to approve or correct.
Letting the model call functions that you define.
Letting the model write code and run it in a sandbox.
A standard way to connect models to tools and data.
Letting the model operate a screen, a mouse and a keyboard.
A model that plans, acts and checks its own work in a loop.
Everything around the model in an agent: the loop, tools, context handling, permissions, caps and sandbox.
An agent that runs its own searches until it has an answer.
Agents that read, write, run and test code.
Reusable instructions that an agent loads when it needs them.
Agents you talk to in real time.
A lead agent splits the task and hands parts to other agents.
Describing a team of agents and how work passes between them.
Agents that check, or argue with, each other's work.
Tasks that run for hours or days.
Agents that resume work across sessions, schedules, and events.
Large groups of agents with roles and shared goals.
Models that control robots and other machines.
Measuring whether a change made the results better.
The tools that run test sets and graders for you, and what to check before trusting their numbers.
Fine-tuning, distillation, synthetic data and automated prompt tuning.
Training a model further on your own examples, in full or with small adapters such as LoRA.
Training a smaller model to reproduce what a larger one does on your task.
Using a model to write training or test examples, and checking them before they are used.
Letting a program search for better prompts against a test set.
Prompt injection, permissions, data handling and audit.
Checks on what goes into a model and what comes out, and the limits of those checks.
Attacking your own system on purpose, before someone else does, and turning what you find into tests.
Cost, speed, monitoring and running models on your own hardware.
Recording what each run did, so a bad result can be traced to the step that caused it.
One entry point in front of several model providers, for keys, routing, limits, fallback and logs.
Spending fewer tokens and less time for the same result: caching, batching, smaller models, shorter context.
Running open-weight models on your own hardware: what fits, quantization, and what you give up.
How to brief a model, review its work and decide what to hand over.
Saying what you want clearly enough that the model does not have to guess.
Checking work you did not do yourself before it goes anywhere.
Deciding which parts of a task to hand to a model and which to keep.
Learning, from results over time, how much to rely on a model without checking.
Uses RAG, structured output and an eval set. Level 2 is enough because a single search answers most questions.
Sorts mail into fixed categories and produces structured output. A person approves anything that gets sent. The categories are known in advance, so an agent is not needed.
Uses agentic RAG to find sources and a fixed check on every claim against the section it cites. It needs level 5 for the searching; the checking is level 3.
A coding agent that reads, edits, runs and tests code in your repository, using skills for repeated tasks and a safety review before anything ships.
Reads the image or PDF, fills a fixed schema, and saves the record once a person confirms it.
Transcribes a voice note, splits it into steps, and turns each step into a structured entry that an eval set checks for accuracy.
Routes an incoming ticket, searches the documentation for an answer, calls a tool when an action is needed, and hands off to a person when it is unsure.
Runs on a timer, diffs a set of public pages in code, and asks a model one question about each change. Level 3: the schedule and the checkpoint are infrastructure, not agency.
A single agent writes and runs code against a dataset, one question at a time, to answer questions a fixed query could not anticipate.
One prompt drafts a piece of writing and another checks it against a rubric, repeating until the draft passes.
Several always-on agents split personal tasks among themselves, sharing memory and staying inside the same safety rules.
Turns a plain-language description of work done into a structured log entry, saved with a tool call and linked to the equipment it concerns through a small knowledge graph.
Organize renewal dates, file names, category totals, and reminders with ordinary code. No model is needed; extracting information from scanned bills is a separate task.
Turn a transcript into decisions, owners, and open questions in one model call. Someone who attended reviews the draft before it is shared.
Extract invoice fields, then use code to match purchase orders and compare amounts. Differences go to a person; the model never decides whether the totals reconcile.
Check an agreement against a fixed checklist, with cited clauses for each finding. Merge the findings for a person to review.
Turn an incident write-up into a timeline and repeatable steps. Check owners and success criteria, then ask the incident lead to approve it.
Split a script into scenes and shots, then check that every line is covered and every shot has a source. A person reviews the plan; drawing frames is a separate task.
Checking what is available, what is open and what connects takes a different number of steps every time, which is what level 5 is for. Read-only lookups run unattended; anything that spends money stops for a person, with the price and the cancellation terms in front of them.
Two independent reviewers apply the same rubric. Disagreements go to the teacher rather than being averaged away.
Code detects new records from fixed sources. One model call summarizes each new title and abstract; code attaches the original citation. It does not follow references or choose new searches.
Code assembles the weekly figures; one model call drafts the report. Checks flag unsupported numbers and missing required facts, then a person reviews and sends it.
Keep a shared tracker current through source comparisons and a review queue. Model proposals and changes to human-written fields need approval; missing evidence is flagged.
Use code to calculate limits, yield, process capability, and trends across lots and fixtures. The pass/fail decision stays deterministic; no model is involved.
Sweep prototype boards across line, load, and temperature. Code calculates margins, uncertainty, and guardbanded verdicts; no model is needed.
Turn computed measurements and notebook notes into a report. Code owns the figures, the model writes the prose, and a person checks the finished draft.
Retrieval over the documents an engineer already has, answered with citations that can be checked. The case that matters is a change notice contradicting the datasheet on one number, where the right answer depends on the board revision. Level 2 is enough because one search finds the passage.
Extract specification rows from a manual, validate their structure, and calculate uncertainty in code. A person verifies ranges, intervals, and conditions against the source.
Failing measurements and free-text operator notes are sorted into the causes the failure analysis guide already lists, then routed. A person confirms before anything is scrapped or reworked. The categories are known in advance, so this is classification into fixed classes and not an agent.
A bill of materials and a netlist summary are checked rule by rule against the written design rules. One pass drafts findings, a second checks each finding against the rule text it cites and drops the ones that cite nothing. Level 3, because code decides every step and the rules do not change between boards. This is the rule check that happens before a review meeting, not the design review report itself: for the report, and the characterization data behind it, see the two recipes this page links in its first paragraph.
A fixed chain: read the requirements, propose a test for each, build the traceability table, then check that every requirement has a test and every test names a requirement. A person approves before any of it is adopted. The order of the steps is known in advance, which is what keeps this at level 3.
The model drafts commands from the manual for that instrument; code checks every one against the documented command set, runs the script on the simulated instrument, and feeds the errors back for another pass. A person bench-checks before it drives real hardware, and every set point goes through a code-side envelope.
Starts where the dashboard stopped: limits, yield and Cpk are already charted and did not answer the question. The model writes analysis code that runs in a sandbox over the CSV, and a person reads the code as well as the answer. Includes the trap of a column in millivolts under a header that says volts.
An agent with read-only tools, instrument queries, the test log and the datasheet, works a low output down to a cause and proposes the next measurement. Queries run unattended; anything that sets a voltage, a current limit or an output goes through the envelope and a person. Level 5 because each measurement depends on the last.
Decoded into Agentic RAG and deep research, Parallel calls, Review and debate, Lead agent and workers.
Decoded into The agent harness, Single agent, Coding agents, Skills, Lead agent and workers, Safety, privacy and governance.
Decoded into Always-on assistants, Long-running tasks, Computer and browser use, Human approval.
Decoded into Retrieval-augmented generation (RAG), Context engineering, Routing.
Decoded into Retrieval-augmented generation (RAG), Embeddings and search, Knowledge graphs and GraphRAG, Safety, privacy and governance.
Decoded into Computer and browser use, Single agent, Human approval, Safety, privacy and governance.
Two uses of graphs that are often confused: graphs that connect information, and graphs that connect work.
The same question asked at every level: which part of this does a person still decide? The answer moves from reading each result to setting the limits a run happens inside.
How you tell whether it worked, from a person reading one answer to a scored set and a recorded trace. The check changes shape at every level; the question does not.
One question followed across the site: what is in front of the model this turn, and who decided to put it there. The answers run from a written instruction to a note a session leaves for the next one.
Use ordinary code, search, forms, or a task-specific statistical model when they solve the problem. No generative model is required; classical machine learning can belong here too.
Give the model instructions and receive a response. A conversation repeats this interaction, with a person directing each turn. Prompting, structured output, reasoning, and multimodal inputs can all fit this pattern.
Add relevant documents, retrieved passages, or stored information to the current request. This supplies context beyond the model’s training without retraining it. Missing, stale, or misleading material can still produce a poor answer.
Connect model calls through predefined steps, branches, checks, and retries. A model can classify an input or evaluate a result to route the workflow; software still defines the available paths.
The model can request a search, calculation, code execution, or an action in another application. Software enforces permissions, performs the action, and returns the result. Tool use alone does not create an ongoing agent loop.
The model uses the goal and observed results to choose an action, revise its approach, or finish. Software executes tools and enforces permissions, approvals, and stopping limits. A run can stop because it is complete, blocked, or out of budget.
Agents divide, coordinate, or review work across separate contexts. A coordinator can combine their findings, and the agents may use the same model or different models. Coordination adds overhead, and separate reviewers can still make correlated mistakes.
Saved state, schedules, and events let an agent start or resume work without a fresh chat message each time. The model need not run continuously, and a dedicated computer or agent team is optional. Permissions, human approvals, monitoring, and stop controls still apply.
Five topics cut across all the levels. Each has its own set of pages.
Context, workflows, reasoning, agents, tools, persistence, and evaluation tradeoffs
Retrieve the relevant policy, answer each part of the question, and distinguish an unknown fact from a retrieval miss.
Extract a useful JSON record, preserve missing fields, and catch a total that does not reconcile.
Extract evidence into a checked table, then draft an update from that table in a fixed two-call workflow.
Draft a calendar change, bind review to the exact proposal, and detect stale or repeated approvals.
Let a model choose read-only diagnostic tools, then require an evidence-backed handoff within six calls.
Process a stock event, save a local outbox record, and prove that replaying the same event does not create another alert.
Human observation guide and results template. Sessions pending.
Lead agent and workers: brief writing, parallel receiving agents, independent reviewers, and coordinated revision. Includes supporting evaluation methods.
98 scripted examples across everyday life, engineering, and business.
The evolution of conversational AI through context, workflows, tools, agent loops, optional teams, and persistent operation. Architectural changes explained with diagrams.
Describe your task and desired automation for your own AI.
Agent instructions, workflows, audits, tool specifications, checks and handoffs.
Describe how an agent should work in your project, then let it verify the details against your actual files.
Define the complete experience from trigger to delivered result, including your role and exceptions.
Turn “it works” into observable checks for quality, completion, and the effort still required from you.
Find where an existing process loses time or quality before deciding what to replace.
Describe a reusable tool an agent can build or adapt, including its contract and how to verify it.
Give the next person or agent a factual starting point, without confusing plans with completed work.
Every technique as a node, laid out by level, with the taxonomy's requires, upgrades-to, combines-with and alternative-to relations as edges.
When each level reached the public, marked at the launch of the product that brought it there, with the paper that first described it behind each mark.
Answer a few questions about your own job and explore a candidate level, then check automation and human effort.
Common tasks and the levels they need, with the use case by technique matrix underneath.
Products you have used, taken apart into the techniques they are built from.
Every developer, model, product and tool this site names, with the technique each one demonstrates and the date it was checked.
Every named failure mode across every technique, grouped by level, each with how to notice it and how to test for it.
Every term a newcomer meets on this site, defined from the page that explains it.
How a level is defined, how a page is written and reviewed, and what the numbers on this site do and do not mean.
The kinds of job people bring to a model, sorted by the shape of the work rather than its subject, with what moves each one to a lower or a higher level.
A dated record of what changed on this site and why it matters to a reader, newest first, with an Atom feed.
The procedure this site hands a reader's own AI assistant: what to ask, how to choose a design that fits the desired outcome, automation, and user effort, and what not to claim.
A model put in charge of a loop rather than one choice: it acts, reads the result, and decides what to do next and when to stop, instead of following steps your code chose in advance.
A team of agents described as a graph: nodes are agents rather than fixed steps, and the output of one node (usually called the supervisor) picks which agent runs next from a fixed list of names.
The repeating cycle behind every agent: the model proposes one action from what it currently sees, your code carries it out, and the result goes back to the model, until the model itself decides to stop.
A loose umbrella for systems where the model, not your code, chooses each step and decides when the task is finished. On this site that begins at level 5; below it a person or a program picks the steps and the model only fills them in.
Retrieval where the model, not your code, decides how many times to search, what to search for next, and when it has read enough to answer, instead of searching once and answering once.
One entry point every call to a model provider goes through instead of calling each provider directly, so provider keys, routing and fallback, per-caller budgets, caching, logging and policy checks sit in one place rather than in every application that calls a model.
An agent built around a computer of its own that keeps running between the moments a person talks to it, deciding on each scheduled check whether anything needs doing and acting under a fixed approval policy.
A fixed, code-side table that sorts every action a model proposes into one of a few classes (run it automatically, queue it for a person, or never run it), regardless of what the model asked for.
A record of what an agent did and why, kept independent of the agent itself, that stands in for a person who was not there to catch a problem as it happened.
Searching for a better prompt against a measured score instead of a person hand-editing the wording, tuning instructions, examples or weights the way a training run tunes a model.
How much of a task a system settles for itself. On this site it is not one quality but a question asked level by level: who decides the next step, and what a person is still holding at that level.
A way to submit many model requests at once for processing that finishes within a day rather than instantly, usually billed at roughly half the price of an ordinary synchronous call.
Generating several candidate answers to the same question and keeping the one a checker judges best, rather than trusting whatever the first attempt produces.
Writing down everything a model needs in order to act without guessing: the goal, context it lacks, constraints, what a finished result looks like, what to do when unsure, and the output format.
Keeping how much you rely on a model without checking in line with how often it has actually been right on tasks like the one in front of you, tracked per task type, not as one overall impression.
A hard limit your code enforces on a loop, such as a maximum number of steps or tokens, that forces a stop the model cannot override or even see.
A saved snapshot of a workflow's shared state, written after a step, so a crashed or interrupted run can resume from that point instead of starting over from the beginning.
A passage a document is cut into before indexing, small enough to embed and retrieve on its own, ideally holding one complete idea rather than splitting a sentence or table row in half.
A pointer from part of an answer back to the specific source passage it came from, so a reader can check whether the source actually supports what the answer claims.
The share of questions where every source a grading rule expects was actually cited in the model's answer, used to score how well a retrieval step is working.
Letting the model write a small program instead of choosing among named tools, then running that program in a sandbox your code controls rather than trusting or interpreting it directly.
A single agent whose tools read files, edit them, run commands and run tests, looping on a propose-edit-run-test cycle until its own tests pass or a cap ends the run.
Replacing an aging context window with a short written summary once it nears its limit, so a session can carry forward what mattered without keeping the full history.
Letting the model operate a real screen: it looks at a screenshot, picks one action such as a click or a keystroke, your code carries it out, and a new screenshot goes back.
A record of where a piece of media came from, carried with the file as signed assertions, specified by C2PA rather than by one maker. It says whether that history validates and is free from tampering, not whether the history is good or bad.
Deciding what goes into a model's request (which instructions, examples, documents and history) and in what order, since the model only knows what it was trained on and what the request contains.
The amount of text a model can read in one request; material that does not comfortably fit has to be trimmed, retrieved, or summarized before the model ever sees the question.
How closely two vectors point in the same direction, used to measure how related two pieces of text are once both have been turned into embeddings.
Secure storage that holds a password, key or payment method on an agent's behalf, so the agent can use it without ever seeing the raw value itself.
Deciding which parts of a task to hand to a model and which to keep, based on what a wrong answer would cost, how checkable the result is, and how reversible the action is.
Training a smaller model to imitate a larger model's outputs on a given task, so the smaller one can stand in for the larger one on that same narrow job.
A list of floating-point numbers standing in for a piece of text's meaning, positioned so texts with similar meaning get vectors that point in similar directions.
Running the same fixed set of questions against a system before and after a change, graded the same way both times, so a claim that the change helped can be checked instead of assumed.
A tool that holds a dataset of examples, runs a program or a prompt against every one of them, grades each result and lets two runs be compared, instead of a team building that machinery from scratch. Some also record a trace of what happened inside each run.
One or more worked examples included in a prompt to show the model a format or pattern rather than only describing it in words.
Training a model further on your own examples so its behavior on that kind of task becomes more consistent, without repeating the same instructions in every request.
Giving the model a fixed list of actions your code defined, each with a name, a description and an argument schema, and letting it choose whether to use one, which one, and what arguments to send.
A fixed list of questions with a known right answer, or a rubric for judging one, run the same way before and after a change so two runs can be fairly compared.
The thing that scores an answer against a golden set, either by matching it exactly against a pattern or by having another model read it against a rubric, which is itself a judgment call worth checking by hand.
One phrase for two different techniques that happen to share a data structure. A knowledge graph connects information: entities and the relationships between them. A workflow graph or an agent graph connects work: steps, and who or what picks the next one.
Tying a claim in an answer to a specific source that can actually be checked, the way RAG grounds an answer in retrieved documents instead of whatever the model remembers from training.
A check on what goes into a model or what comes out: an input filter, an output validator, a separate classifier trained to judge safety, a schema check. It is probabilistic and can be wrong in both directions, so it is never the control that holds; a code check that tests a specific fact is.
A fluent, confident answer that is not actually true or not supported by any real source; research argues this happens because training and grading reward a plausible guess over admitting uncertainty.
The edge in an agent graph where one agent's output hands control to another named agent, chosen by a model call rather than a rule your code wrote in advance.
Everything around the model in an agent: the loop that calls it, the tool definitions it is shown and the code that runs them, what goes into its next request, whether an action needs approval, the caps on steps and tokens, the sandbox, and what gets logged. None of it is the model.
A pause your code inserts before something costly, irreversible or too uncertain to ship, handing the decision to a person instead of letting the run continue on its own.
Running a keyword search and a meaning-based search over the same documents and merging the two result lists, so an exact identifier and a paraphrased question can both be found.
The set of embeddings, or the keyword structure, built from a document set in advance so a later question can be compared against it and ranked, without re-reading every document.
The date after which a model's training data stops, so anything that changed after that date is not something the model actually knows, however confidently it answers.
Facts stored as entities and the relationships between them, so a chain of hops can join facts across documents instead of needing one passage to state the whole answer.
How long a request takes to complete, tracked separately from cost; caching, batching and routing to a smaller model are among the main ways to bring it down at volume.
Giving a system only the access its task actually needs, so an action it was never granted is one no instruction, however cleverly worded, can talk it into taking.
One of the eight levels this site sorts a technique into. A new level starts where the answer to “who decides the next step” changes: nobody, you, your code, the model for one action, the model for every step, several models, the models including when to start.
A model run on your own hardware instead of called over an API, trading a per-call bill for hardware you own and buy, and for models small enough to fit on it.
Work that starts on a schedule or an event and continues across many separate sessions until its queue or goal is finished or a person steps in, with no single session seeing the one before it directly.
Designing an agent's loop on purpose (what starts it, what it repeats, what stops it) rather than only writing its prompt. Anthropic's Claude Code team defines loops as agents repeating cycles of work until a stop condition is met.
A small, trainable add-on layered onto a model's frozen weights instead of retraining the whole model, cutting the parameters and memory a fine-tuning run needs by orders of magnitude.
A standard way for an application to connect to servers that expose tools, resources and prompts to a model, instead of a developer wiring each integration by hand.
Keeping information from one conversation to the next by writing facts down somewhere and reading them back into a later, otherwise unrelated conversation, since a chat has no memory of its own.
The part of a product that writes facts down after one conversation and reads them back into a later one. Two different things get called this: a summary, cheap to reread but lossy, and a record kept on its own and searched on demand.
A step in a recorded run where the model's own output chose what happened next (which tool, which query, whether to stop), as opposed to a step your code decided regardless of what the model said.
Several agents working on a task, using the same underlying model or different models: a lead agent splitting the work among others, agents handing work to each other across a graph, or two agents checking each other's output. That is level 6 on this site.
A model that takes or produces more than text (images, audio, video and documents) inside the same one-call request and response shape as an ordinary chat message.
Recording what each run did in enough detail that a bad result can be traced back to the step that caused it: which passages a retrieval step picked, which tool the model called and with what arguments, which branch a workflow took, and what each step spent.
A lead model that reads a task, decides how to split it, hands each piece to a worker, and combines what comes back, rather than doing the whole task itself.
Several standing agents with distinct roles and a shared goal, where the roster itself keeps running and changing over time rather than being assembled fresh for one job and torn down after.
Running more than one model call at the same time instead of one after another, then combining the results in code, either splitting one task into independent parts or running the same task several times to vote.
Loading only a skill's short name and description into context up front, and its full instructions only once the model actually decides to use it, so unused skills cost almost nothing.
The request text sent to a model in one call: instructions, examples and the question, everything the model sees that is not already baked into its training.
Reusing a matching, unchanged prefix of a request across calls at a reduced billing rate, which is why makers recommend putting content that never changes first and content that changes every call last.
Splitting one task into a fixed sequence of steps and handing each step's output to the next, with your code deciding how many steps there are and what gate sits between them.
Writing the request itself well (instructions, examples, an assigned role, a required format, sometimes asking the model to reason first) to get a more consistent result from a single call.
Text written to look like an instruction, hidden in the user's message or in retrieved or tool-returned content, that tries to redirect what the model does instead of answering the actual question.
Storing a model's weights at lower precision so a large model fits smaller hardware and may run faster. llama.cpp's own documentation says it shrinks the model and can speed up inference, and that it may cost some accuracy. Name the exact level you ran, not just "4-bit".
A persisted list of pending work items a long-running task drains one at a time across separate sessions, surviving a crash because the queue is checkpointed after every item, not held only in memory.
Searching your own documents for the passages closest to a question, putting those passages in the prompt, and asking the model to answer once using only what it was given.
A label rather than a defined technique: different people use it for different bundles of retrieval improvements: better chunking, reranking, hybrid search, retrieval the model itself drives. This site has no page under that name; the pages on RAG, embeddings and search, and agentic RAG cover the substance.
How much a model is allowed to work through a problem in its own words before answering; makers recommend more of it for math, debugging and planning and less for simple lookups.
A model trained, usually with reinforcement learning, to write out a long chain of thought before it answers, and to notice and fix its own mistakes along the way. OpenAI's o1 (September 2024) was the first sold as one.
The operation that ranks stored facts or documents against a new question and returns the closest matches, the same mechanism whether the store is a document corpus or a small set of personal facts.
Attacking your own system on purpose, before someone else does it without permission, across the system around the model and not only the model itself. What matters is not the report but the change it forces, most durably a test that keeps failing until the finding is fixed.
A second pass over a first, larger cut of retrieved candidates, scored by a model trained specifically to sort results by relevance to the original query.
Searching an index of document chunks for the ones closest to a question, and keeping a fixed number of the best matches to hand to the model.
Checking work you did not do yourself before it goes anywhere: pulling out specific claims and checking them against their source, then checking what is missing, then judging the result.
A separate agent, with its own context and often its own retrieval, checks or argues with another agent's work and decides whether to accept it, instead of one fixed, code-owned test.
Looking at an input, deciding which of several fixed kinds it is, and sending it down the handler built for that kind, with a real fallback for whatever fits none of them.
A written checklist a grader model reads an answer against when the answer is too open-ended to match against a fixed pattern, used to turn a judgment call into a repeatable score.
An isolated environment with no network access and fixed resource limits, where code the model wrote is actually run, so what the model produces is data your code hands to an interpreter, never code it trusts directly.
A fixed shape for a reply (named fields with defined types) that a model's output is constrained to match, so downstream code can parse it without guessing at its structure.
Asking the same question several times as independent calls and returning whichever answer the largest share of the samples agree on, instead of trusting a single attempt.
Instructions an agent keeps on the shelf until it decides it needs them, loaded into context only when triggered, instead of text repeated into every single turn like a system prompt.
The current MCP specification, revision 2026-07-28, removed the initialize handshake and protocol-level sessions. Every request now carries its own protocol version and capabilities, and a server that needs state across calls returns an explicit, server-minted handle the client passes back as an ordinary tool argument.
The test a model-driven loop uses to decide it is finished and should give a final answer rather than take another action; a cap can force a stop the model never actually reaches.
A reply constrained to come back in a fixed shape, such as a JSON object with named fields, instead of a paragraph your code has to parse by guessing.
Training examples generated with a model rather than collected from real use, for a later fine-tuning or evaluation run.
TypeSafe AI's name for a model that generates no text and returns a typed decision with a probability in one fast pass, for use inside software. Jev (September 2026) is the first. The name borrows psychology's fast System 1, as against slow, deliberate System 2.
The instructions sent to a model at the start of a request, separate from what the user or the retrieved content says, setting how it should behave for that call.
Computation spent while answering, as opposed to while training. Reasoning models made it a dial: more thinking time, better answers on hard problems, at the cost of seconds and output tokens.
The unit a model's input and output are measured and billed in; every cost strip on this site counts tokens in and tokens out for the run it illustrates.
A recorded, step-by-step account of a run (every model call, every tool call, and whether each step was decided by code or by the model) used to explain how an answer was actually produced.
The path a run took to reach its answer: which tools were called, in what order, with what arguments. Scoring it is a separate measurement from scoring the final answer, and it gives partial credit for the steps a run got right.
A fixed-length list of numbers representing a piece of text, an embedding, positioned in space so that texts with related meaning end up close together.
A database built to hold embeddings and rank them by similarity to a query vector: pgvector, Pinecone, Weaviate and Qdrant are four. It stores the index a retrieval step searches; the embeddings it stores come from an embedding model, not from the database.
A model that converts vision and language input directly into motor control, letting a robot take a physical action rather than only produce text.
A single agent you talk to instead of type to, in real time, that additionally has to decide when a caller has stopped talking and what to do if they interrupt.
A loop where one prompt writes a draft and a separate prompt checks it against one specific, testable criterion, revising and rechecking until it passes or a fixed revision cap is reached.
Weizenbaum's paper describes ELIZA, a program that scans input for keywords and applies decomposition and reassembly rules attached to them; the paper says "Keywords and their associated transformation rules constitute the SCRIPT for a particular class of conversation." Code picks the reply, with no model involved.
Elastic's own history post places Elasticsearch's first release, which it says "happened to be 0.4.0", in February 2010: keyword search an ordinary developer could run without writing a search engine.
scikit-learn's own About page says these four "took leadership of the project and made the first public release, February the 1st 2010", packaging classical machine learning algorithms a developer could call without implementing them.
The paper describes a scalable tree boosting system "used widely by data scientists to achieve state-of-the-art results". Classical machine learning, no language model anywhere in it, still advancing in the middle of the deep-learning decade.
Google researchers introduced the Transformer, an architecture built on attention alone, dispensing with the recurrence and convolutions earlier sequence models relied on. It is the architecture under the one-call language models the rest of this level names.
GitHub's record of huggingface/transformers gives a creation date of October 29, 2018. The repository describes itself as "the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training": the open library a developer could load a pretrained language model with, years before anyone could buy a chat product.
OpenAI's post on GPT-2 reports the model doing reading comprehension, translation and summarization "without any fine-tuning of our models, simply by prompting the trained model in the right way". That is level 1 stated plainly: write the request, read the response, no training step in between.
A text adventure built on GPT-2: the player types any action in plain English and the model writes what happens next. Its maker's own post is dated Thursday, December 5, 2019. One request, one response, open to anyone, three years before ChatGPT. At first it ran from a shared Google notebook, not an app.
OpenAI's paper describes "an autoregressive language model with 175 billion parameters" applied "without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction". That is the one-call pattern that prompt engineering works within.
OpenAI's announcement of June 11, 2020 says “We’re releasing an API for accessing new AI models developed by OpenAI”, providing “a general-purpose ‘text in, text out’ interface”, and that “Today the API runs models with weights from the GPT-3 family”. Access was not open: “we are launching today in a private beta rather than general availability”, with a waitlist.
The paper reports that "generating a chain of thought -- a series of intermediate reasoning steps -- significantly improves the ability of large language models to perform complex reasoning", and that it "improves performance on a range of arithmetic, commonsense, and symbolic reasoning tasks". Nothing about the model changes; only the prompt does.
OpenAI fine-tuned GPT-3 on human demonstrations and human rankings of its outputs. The paper reports that "outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters", in human evaluations on OpenAI's own prompt distribution.
OpenAI's announcement, dated November 30, 2022, says “We’ve trained a model called ChatGPT which interacts in a conversational way” and “During the research preview, usage of ChatGPT is free. Try it now at chat.openai.com.” A chat product anyone could open in a browser, with no API key, no code and no invitation.
OpenAI's technical report describes "a large-scale, multimodal model which can accept image and text inputs and produce text outputs", and reports that it "exhibits human-level performance on various professional and academic benchmarks". That is OpenAI's own evaluation of its own model.
OpenAI's API changelog entry for May 13, 2024 reads: "Released GPT-4o in the API. GPT-4o is our fastest and most affordable flagship model." The changelog does not describe how its modalities are combined; this site cites it only for the release date and OpenAI's own description.
OpenAI's API changelog entry for August 6, 2024 reads: "Launched Structured Outputs—model outputs now reliably adhere to developer supplied JSON Schemas."
OpenAI's post of September 12, 2024 opens: "We are introducing OpenAI o1, a new large language model trained with reinforcement learning to perform complex reasoning. o1 thinks before it answers—it can produce a long internal chain of thought before responding to the user." It adds that the model "learns to recognize and correct its mistakes" and that its performance improves "with more time spent thinking (test-time compute)". The extra work happens inside one request and response, which is why this sits at level 1 and not higher. OpenAI's API changelog for the same day records the release of o1-preview and o1-mini.
The paper reports that "the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories", and that the training brings out "advanced reasoning patterns, such as self-reflection, verification, and dynamic strategy adaptation." A second lab, in the open, reaching what o1 had shown four months earlier.
Google's Gemini API changelog entry for February 19, 2026 reads: "Released Gemini 3.1 Pro Preview, our latest iteration in the new Gemini 3 series family."
OpenAI's API changelog entry for September 3, 2026 reads: "Released GPT-6 Astra, our most capable model, built for the hardest end-to-end work." That is OpenAI's own description of its own model.
DeepSeek's release note of September 10, 2026 says "V4.1-Flash is now live on the DeepSeek API with native multimodal support", describing a mixture-of-experts model with 552B parameters and 8B active for input, 16B for output: a frontier-class open-weight release, one call at a time.
TypeSafe's founder, Diogo Almeida, announces "a new class of frontier models built to make fast, structured decisions that software can use directly": "unstructured state in, typed probabilistic decisions out." Jev writes no text. Its possible outputs are defined in advance, every answer carries a probability, and all of it comes back in one parallel pass, not word by word. TypeSafe quotes "70ms-500ms" a call and "$0.042 / MTok" of input. Its own list of uses is this site's level 3: "classify, route, score, extract, or branch where hand-written logic is too brittle."
GitHub's record of facebookresearch/faiss gives a creation date of February 7, 2017 and describes it as "A library for efficient similarity search and clustering of dense vectors". That is the open library a developer could build vector search over their own documents with, and it is still the substrate under most retrieval code.
The paper introduces RAG, in which "the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index of Wikipedia", "accessed with a pre-trained neural retriever": the model answers from passages retrieved for the question rather than only from its own weights.
OpenAI's post of January 25, 2022 says “We are introducing embeddings, a new endpoint in the OpenAI API that makes it easy to perform natural language and code tasks like semantic search, clustering, topic modeling, and classification”, and that “Embeddings that are numerically similar are also semantically similar.” Vectors as a service: a developer no longer had to train or host an embedding model.
GitHub's record of the repository now published as run-llama/llama_index gives a creation date of November 2, 2022, and the first release of its package on PyPI, then named gpt-index, is dated November 22, 2022 with the summary “Building an index of GPT summaries.” The first widely used open library built for this level's job: index a set of documents, retrieve from it, and hand what comes back to a language model.
Neeva's post, bylined "The Neeva Team on 01/06/23", introduces a written answer at the top of its search results with citations embedded in the text, open at once to account holders in the United States. A month before the new Bing. Neeva closed its search engine later in 2023.
Microsoft's announcement says "Bing reviews results from across the web to find and summarize the answer you're looking for" and that "The new Bing also cites all its sources". On this date it was a limited preview on desktop with a waitlist, so it is not the level's availability date.
Microsoft's announcement says "the new Bing is now in Open Preview and no longer has a waitlist": anyone with a Microsoft account could now ask a question and get an answer written from pages retrieved for it, with links to those pages.
Anthropic's announcement says: "We've expanded Claude's context window from 9K to 100K tokens, corresponding to around 75,000 words!" Enough room to paste the material in rather than retrieve from it.
OpenAI's API changelog entry for January 25, 2024 reads: "Released embedding V3 models and an updated GPT-4 Turbo preview". The changelog entry itself says nothing further about the models; this site cites it for the date only.
Microsoft Research's post describes an approach in which "The LLM processes the entire private dataset, creating references to all entities and relationships within the source data, which are then used to create an LLM-generated knowledge graph", which is then clustered and pre-summarized so questions can be answered across many documents rather than from the nearest few passages.
OpenAI's post of February 13, 2024 says “We’re testing the ability for ChatGPT to remember things you discuss to make future chats more helpful”, that a user “can explicitly tell it to remember something, ask it what it remembers, and tell it to forget conversationally or through settings”, and that memory can be turned off entirely.
Google's announcement says "We can now run up to 1 million tokens in production", and that at this date "a limited group of developers and enterprise customers can try it with a context window of up to 1 million tokens". A private preview, not a general release, and the standard window was 128,000 tokens.
OpenAI's API changelog entry for October 1, 2024 reads: "Prompt caching: Discounts and faster processing times on recently seen input tokens." Anthropic's own API release notes record prompt caching leaving beta on the Claude API on December 17, 2024. Reusing a long shared prefix is what makes a large fixed context affordable to send on every request.
Google's announcement says: "We're renaming NotebookLM to Gemini Notebook. It's the same standalone product, now doing more across the Google ecosystem and updated with a secure cloud computer." The product whose whole premise is answering only from the documents you gave it is still the plainest consumer example of level 2.
The paper proposes "Chaining LLM steps together, where the output of one step becomes the input for the next, thus aggregating the gains per step", and is titled for what that buys: transparent and controllable human-AI interaction, because the intermediate results exist as steps a person can see.
PyPI's release history for the langchain package shows version 0.0.1 uploaded on October 25, 2022, with 0.0.2 the next day: the library a developer could install instead of writing chaining and model-swapping plumbing themselves.
Zapier's own OpenAI integration page offers an action that “Sends a prompt to OpenAI and generate a response” inside a multi-step Zap, and says “Zapier lets you connect OpenAI with thousands of the most popular apps, so you can automate your work and have more time for what matters most—no code required.” The flow decides what runs next; one of its steps calls a model. The page carries no launch date, so the date marked here is the earliest capture of it this site could read: the feature was live by then, and may well have shipped earlier.
Microsoft's announcement says Copilot "is more than OpenAI’s ChatGPT embedded into Microsoft 365. It’s a sophisticated processing and orchestration engine working behind the scenes to combine the power of LLMs, including GPT-4, with the Microsoft 365 apps and your business data in the Microsoft Graph". Software runs the steps (fetch the person's files and mail, build the prompt, call the model, check the result, write into Word or Outlook) and the model fills them in. Shown that day: a first draft in Word from your own files, a deck in PowerPoint from a document, trend analysis in Excel, thread summaries and draft replies in Outlook, and live meeting summaries in Teams.
Microsoft's post of November 1, 2023 opens: "Starting today, Microsoft 365 Copilot is generally available for enterprise customers worldwide."
PyPI shows version 0.0.1 of semantic-router uploaded on November 9, 2023. Its project page calls it "a superfast decision-making layer for your LLMs and agents" that routes requests "using semantic meaning" rather than waiting on a model's generation: code picks the branch.
Microsoft's Power Platform blog says "GPT Prompts with Prompt Builder, a new feature of AI Builder, is now generally available!" and that it lets a person "add content processing and content generation capabilities to Power Automate". An ordinary customer drops a model call into a flow whose next step is still chosen by the flow, not the model.
LangChain's launch post says "LangGraph is module built on top of LangChain to better enable creation of cyclical graphs, often needed for agent runtimes", and that until then "we've lacked a method for easily introducing cycles into these chains." A developer describes the application as a state machine of nodes and edges.
OpenAI's API changelog records the Batch API's release on April 15, 2024. Sending a set of independent requests together, and collecting them when they finish, is the parallel-calls pattern offered as a product feature rather than assembled by hand.
LangChain's post introduces interrupt, which will "pause execution of the graph, mark the thread you are running as interrupted, and put whatever you passed as an input to interrupt into the persistence layer." The pattern it names first: "Pause the graph before a critical step, such as an API call, to review and approve the action. If the action is rejected, you can prevent the graph from executing the step".
Anthropic's engineering post names and diagrams five workflow patterns (prompt chaining, routing, parallelization, orchestrator-workers and evaluator-optimizer) and advises: "Start with simple prompts, optimize them with comprehensive evaluation, and add multi-step agentic systems only when simpler solutions fall short."
Flowise's sunset notice says "we've decided to wind down our operations for Flowise", with a feature freeze on July 29, 2026 and the repository archived on August 10, 2026, adding that "Flowise source code will still remain on Github and the Apache 2.0 licensed code is yours to keep building on." A visual workflow builder closing while the pattern it built stayed in use.
The paper describes a language model surrounded by tools (a calculator, a currency converter, a database call) and "a router that routes every incoming natural language input to a module that can best respond to the input". The router is itself a small neural network. A model chooses the tool and code runs it: the whole paper is about that one idea, nine months before Toolformer.
The paper describes a model trained in a self-supervised way, from a handful of examples per API, to "decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction": the model choosing the action, not code choosing it for the model.
OpenAI's post of March 23, 2023 says “We’ve implemented initial support for plugins in ChatGPT”, tools that “help ChatGPT access up-to-date information, run computations, or use third-party services”, and that “we're also hosting two plugins ourselves, a web browser and code interpreter”. The model picks which plugin to call mid-conversation. This was an invitation: “Today, we will begin extending plugin alpha access to users and developers from our waitlist.”
OpenAI's ChatGPT release notes carry an entry headed “Web browsing and Plugins are now rolling out in beta (May 12)”, which says “If you are a ChatGPT Plus user, enjoy early access to experimental new features” through a beta panel “which is rolling out to all Plus users over the course of the next week”, and describes plugins as “a new version of ChatGPT that knows when and how to use third-party plugins that you enable”. No waitlist and no invitation: a subscription was enough.
OpenAI's post of June 13, 2023 says “Developers can now describe functions to gpt-4-0613 and gpt-3.5-turbo-0613, and have the model intelligently choose to output a JSON object containing arguments to call those functions”, through “new API parameters in our /v1/chat/completions endpoint, functions and function_call”. The model chooses the call; the developer's code makes it.
Google's announcement says Extensions let Bard "find and show you relevant information from the Google tools you use every day — like Gmail, Docs, Drive, Google Maps, YouTube, and Google Flights and hotels — even when the information you need is across multiple apps and services." The model decides which of those to reach for mid-conversation; Google's code makes the call. The product is now Gemini's connected apps.
Anthropic's announcement describes Claude working a computer "by looking at a screen, moving a cursor, clicking buttons, and typing text", and says of the capability: "At this stage, it is still experimental—at times cumbersome and error-prone." It adds that "we encourage developers to begin exploration with low-risk tasks."
Anthropic's announcement describes a tool that lets Claude "write and run JavaScript code directly in Claude.ai" to "process data, conduct analysis, and produce real-time insights", available to "all Claude.ai users in feature preview": code execution reaching people who write none.
Anthropic's announcement says "The Model Context Protocol is an open standard that enables developers to build secure, two-way connections between their data sources and AI-powered tools", and that it provides "a universal, open standard for connecting AI systems with data sources, replacing fragmented integrations with a single protocol."
OpenAI's API changelog for March 11, 2025 records "the Responses API, a new API for creating and using agents and tools" and, the same day, "a set of built-in tools" for it: web search, file search and computer use. Remote MCP servers and a code interpreter were added on May 20, 2025.
The paper interleaves reasoning traces with actions, so that "reasoning traces help the model induce, track, and update action plans as well as handle exceptions, while actions allow it to interface with external sources, such as knowledge bases or environments, to gather additional information." The model reads each result and picks the next move, which is what makes the loop the model's rather than the code's.
GitHub's record of the Significant-Gravitas/AutoGPT repository gives a creation date of March 16, 2023. The project describes itself as "The open-source platform for AI agents" and says it "lets you build, deploy, and run AI agents that carry out complete workflows": a goal-seeking loop a developer could run instead of writing one.
A free web page, live by April 9, 2023: "Assemble, configure, and deploy autonomous AI Agents in your browser. Create an agent by adding a name / goal, and hitting deploy!" It wrote itself a task list, worked through it and added tasks from the results. Part of the AutoGPT wave of spring 2023, which put a looping agent in front of anyone with a browser, and which few people remember as a product.
Cognition's announcement, headed "Introducing Devin, the first AI software engineer", says it equipped Devin "with common developer tools including the shell, code editor, and browser within a sandboxed compute environment—everything a human would need to do their work." On this date it was not something a customer could buy: "Devin is currently in early access as we ramp up capacity."
Replit's post says “Last week, we launched Replit Agent, our AI system that can create and deploy applications”, and that “It configures your development environment, installs dependencies, and executes code”: the model choosing each next step and stopping when the app runs. Replit adds: “The agent is available today in early access to all Replit Core subscribers” and “it should be treated as ‘alpha’ software.” A subscription, not an invitation.
OpenAI's API changelog entry for October 1, 2024 reads: "Realtime API: Build fast speech-to-speech experiences into your applications using a WebSockets interface." The same changelog records it becoming generally available on August 28, 2025.
Cognition's post of December 10, 2024, headed “Devin is now generally available”, says “Today we’re making Devin generally available starting at $500 a month for engineering teams”. The nine months between this and Devin's announcement are the gap between a demonstration and something a customer could buy.
Google's announcement says Deep Research creates "a multi-step research plan for you to either revise or approve", then works "browsing the web the way you do: searching, finding interesting pieces of information and then starting a new search based on what it's learned", ending in "a comprehensive report of the key findings". It launched that day on desktop and mobile web for Gemini Advanced subscribers.
OpenAI's post introduces "A research preview of an agent that can use its own browser to perform tasks for you." The person describes a task; the model looks at screenshots of a browser running in the cloud and clicks and types until the job is done. It hands back when it should: "Operator is trained to proactively ask the user to take over for tasks that require login, payment details, or when solving CAPTCHAs." The post says "Available to Pro users in the U.S."
OpenAI's post of February 2, 2025 describes deep research as “An agent that uses reasoning to synthesize large amounts of online information and complete multi-step research tasks for you”, which “conducts multi-step research on the internet for complex tasks” and returns a report with citations. The same page says: “Available to Pro users today, Plus and Team next.”
Anthropic's announcement says Claude Code, released "as a limited research preview" alongside Claude 3.7 Sonnet, "enables developers to delegate substantial engineering tasks to Claude directly from their terminal", describing it as "an active collaborator that can search and read code, edit files, write and run tests, commit and push code to GitHub, and use command line tools".
Anthropic's announcement says Claude "operates agentically, conducting multiple searches that build on each other while determining exactly what to investigate next", across "both your internal work context and the web". It launched "in early beta for Max, Team, and Enterprise plans in the United States, Japan, and Brazil."
OpenAI's post of July 17, 2025 says “ChatGPT can now do work for you using its own computer, handling complex tasks from start to finish”, with “a visual browser that interacts with the web through a graphical-user interface, a text-based browser for simpler reasoning-based web queries, a terminal, and direct API access”, and that “Starting today, Pro, Plus, and Team users can activate ChatGPT’s new agentic capabilities”. The person still starts every task, which is what keeps this at level 5 rather than level 7.
Anthropic's announcement says "Skills are folders that include instructions, scripts, and resources that Claude can load when needed", and that "Claude will only access a skill when it's relevant to the task at hand": the agent deciding which of its own instructions to read.
Google's post says "we're unifying our efforts into Google Antigravity, our premier agent-first development platform", and that "On June 18, 2026, Gemini CLI and Gemini Code Assist IDE extensions will stop serving requests for Google AI Pro and Ultra, as well as those using it free of charge using Gemini Code Assist for individuals." Enterprise licenses keep Gemini CLI.
Two model agents, one given the role of the person with a task and one the role of the assistant, work the task out between them with no human in the conversation. The abstract proposes "a novel communicative agent framework named role-playing" and reports "comprehensive studies on instruction-following cooperation in multi-agent settings." Eight weeks before the multi-agent debate paper.
The paper has "multiple language model instances propose and debate their individual responses and reasoning processes over multiple rounds to arrive at a common final answer", and reports that this "improves the factual validity of generated content, reducing fallacious answers and hallucinations that contemporary models are prone to".
GitHub's record of the microsoft/autogen repository gives a creation date of August 18, 2023; the first pyautogen release reached PyPI a week later, on August 25, 2023. The README calls it "a framework for creating multi-agent AI applications that can act autonomously or work alongside humans".
A search product that builds a page of results for each question. Its launch post, dated Jun 18, 2024, describes a "multi-agent framework" and a "team of specialized AI agents". It is the earliest product this site found whose own launch page says several agents share one job, ten months before Claude Research. The post never says how the agents divide the work.
Google's announcement says "The A2A protocol will allow AI agents to communicate with each other, securely exchange information, and coordinate actions on top of various enterprise platforms or applications", and says more than 50 technology partners contributed to it.
Anthropic's launch post of April 15, 2025 says Claude “operates agentically, conducting multiple searches that build on each other while determining exactly what to investigate next”, and that “Research is now available in early beta for Max, Team, and Enterprise plans in the United States, Japan, and Brazil.” Anthropic's engineering post two months later says of the same feature: “Our Research system uses a multi-agent architecture with an orchestrator-worker pattern, where a lead agent coordinates the process while delegating to specialized subagents that operate in parallel.” The launch post itself does not say several agents run on one request.
Anthropic's engineering post says of a feature customers were already using: "Our Research system uses a multi-agent architecture with an orchestrator-worker pattern, where a lead agent coordinates the process while delegating to specialized subagents that operate in parallel." The subagents work "with their own context windows" before condensing what they found for the lead agent.
xAI's announcement says: "We have made further progress on parallel test-time compute, which allows Grok to consider multiple hypotheses at once. We call this model Grok 4 Heavy", sold through "a new SuperGrok Heavy tier". Several runs on one question, bought as a tier.
The paper puts "a small town of twenty five agents" in a sandbox. Its architecture, in the paper's words, is one that agents use to "store a complete record of the agent's experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior": agents that decide when to act, not only what to answer.
Physical Intelligence's post says "We've developed a general-purpose robot foundation model that we call π0 (pi-zero)" that "spans images, text, and actions and acquires physical intelligence by training on embodied experience from robots, learning to directly output low-level motor commands", trained on open-source data plus dexterous tasks collected "across 8 distinct robots".
Manus describes itself, on its own site, as building "general AI agents as the Action Engine for life" and as "building the hands for AI to do" rather than the reasoning underneath: an agent a person points at a goal and lets run.
Google DeepMind's announcement introduces "Gemini Robotics, an advanced vision-language-action (VLA) model that was built on Gemini 2.0 with the addition of physical actions as a new output modality for the purpose of directly controlling robots", alongside Gemini Robotics-ER for spatial reasoning. The registry's Gemini Robotics entry is the later Gemini Robotics 2.
GitHub's record of NousResearch/hermes-agent gives a creation date of July 22, 2025 and an MIT license. The README calls it “The self-improving AI agent built by Nous Research” and says “Run it on a $5 VPS, a GPU cluster, or serverless infrastructure that costs nearly nothing when idle”; the project's own site (hermes-agent.nousresearch.com) says “Tasks that need an active agent will not run while it is stopped; hosted agents are managed separately in Nous Portal.” That is standing work that runs because the agent is running, on hardware the developer keeps up.
GitHub's record of openclaw/openclaw gives a creation date of November 24, 2025. Its README says “OpenClaw is an open-source AI assistant that runs on your own computer and meets you in the channels you already use”, with “One Gateway” running it “as a personal assistant on a laptop or as a shared team deployment”, and carries an MIT license badge.
Anthropic's dated release notes record, on January 12, 2026, “Cowork research preview on Claude Desktop (macOS only) for Max plans”, and say Cowork “brings Claude Code’s agentic capabilities to the Claude desktop app for knowledge work beyond coding”. Pro plans followed on January 16, 2026 and the same notes record “Claude Cowork is now generally available on macOS and Windows through the Claude Desktop app” on April 9, 2026.
Physical Intelligence's post calls π0.7 "a steerable generalist model that can perform dexterous tasks across robots, scenes, and skills", and reports "compositional generalization, recombining skills from various tasks to solve new problems", including a robot folding laundry with no laundry-folding data in its training.
Google's announcement introduces "Gemini Spark, a 24/7 personal AI agent that helps you navigate your digital life", "deeply integrated with the Workspace tools you rely on daily, like Gmail, Docs, Slides and more", and says that "because it is a cloud-based agent, Spark keeps working in the background even when you close your laptop or lock your phone."
Google's post of June 30, 2026 says "Gemini Spark for macOS is available in Beta to Google AI Ultra subscribers aged 18 and over, starting in the US." The May announcement had described Spark as "a 24/7 personal AI agent" that takes "recurring tasks or triggers" and "keeps working in the background even when you close your laptop or lock your phone."
Anthropic's dated release notes record, on July 7, 2026: "Cowork runs your sessions remotely (in beta), so your sessions and files are saved to your Claude account and go where you go, on any device. Work continues when you close your laptop, and scheduled tasks run with no device online." This is the entry that makes Cowork an always-on product; the January release ran on the person's own machine.
OpenAI's product page says “Powered by GPT‑5.6, ChatGPT Work brings together context from your team’s tools to turn scattered notes, drafts, and ideas into finished work — and keeps projects moving while you stay in control”, and that it “gathers context, plans the approach, and takes action across your tools, files, and desktop apps”.
xAI's announcement says "Grok Bot is your team of always-on agents" and, of those agents, "They have their own computer, work inside tools and apps like you do, and keep working 24/7." It says Grok Bot "is in beta and available today" to named SuperGrok and Cursor subscriber tiers on desktop and iOS.
Meta's announcement says "Muse is a personal AI agent. It doesn't just answer questions, it actually does the work", and that "A separate Sentinel agent runs on that same machine, kept apart from Muse at the system level. Nothing Muse does reaches the internet unless the Sentinel approves it, and it asks the person for permission when needed." A checking agent shipped inside a consumer product, not only proposed in a paper.
Anthropic's announcement folds a separate always-on product back into the chat app: "hand over a report due at noon, and Claude takes it from there, even after you've closed your laptop", and "You can check progress from your phone on the way to the office." Cowork's capabilities are now "available from any conversation, with the context, skills, and connectors you already have."
The early Perplexity Ask generated answers from retrieved search results with citations. Cofounder Aravind Srinivas retrospectively dates its launch to December 7, 2022 at Stripe Sessions 2024.