Glossary

106 terms, defined from the pages that explain them

Every definition below is written from the technique or topic page named beside it, so the glossary cannot say something that page does not. A definition is one or two plain sentences a newcomer can read. Click a term's name to read the page it comes from.

A

agentLevel 05 · Agent loops

A model put in charge of a loop rather than one choice: it acts, reads the result, and decides what to do next and when to stop, instead of following steps your code chose in advance.

See also: agent loop, cap, stop condition

agent graphLevel 06 · Teams of Agents

A team of agents described as a graph: nodes are agents rather than fixed steps, and the output of one node (usually called the supervisor) picks which agent runs next from a fixed list of names.

See also: handoff, checkpoint, agent loop

agent loopLevel 04 · Tool use

The repeating cycle behind every agent: the model proposes one action from what it currently sees, your code carries it out, and the result goes back to the model, until the model itself decides to stop.

agentic AILevel 05 · Agent loops

A loose umbrella for systems where the model, not your code, chooses each step and decides when the task is finished. On this site that begins at level 5; below it a person or a program picks the steps and the model only fills them in.

See also: agent, agent loop, level

agentic RAGalso called deep researchLevel 05 · Agent loops

Retrieval where the model, not your code, decides how many times to search, what to search for next, and when it has read enough to answer, instead of searching once and answering once.

See also: RAG, agent loop, retrieval

AI gatewayTopics at every level

One entry point every call to a model provider goes through instead of calling each provider directly, so provider keys, routing and fallback, per-caller budgets, caching, logging and policy checks sit in one place rather than in every application that calls a model.

See also: latency, local model

always-on agentalso called always-on assistant, agent teammateLevel 07 · Always-on agents

An agent built around a computer of its own that keeps running between the moments a person talks to it, deciding on each scheduled check whether anything needs doing and acting under a fixed approval policy.

See also: approval policy, audit trail, long-running task

approval policyLevel 07 · Always-on agents

A fixed, code-side table that sorts every action a model proposes into one of a few classes (run it automatically, queue it for a person, or never run it), regardless of what the model asked for.

audit trailTopics at every level

A record of what an agent did and why, kept independent of the agent itself, that stands in for a person who was not there to catch a problem as it happened.

automated prompt optimizationTopics at every level

Searching for a better prompt against a measured score instead of a person hand-editing the wording, tuning instructions, examples or weights the way a training run tunes a model.

autonomyThread

How much of a task a system settles for itself. On this site it is not one quality but a question asked level by level: who decides the next step, and what a person is still holding at that level.

See also: human approval, approval policy, level

B

batch APITopics at every level

A way to submit many model requests at once for processing that finishes within a day rather than instantly, usually billed at roughly half the price of an ordinary synchronous call.

best-of-NLevel 01 · Direct prompting

Generating several candidate answers to the same question and keeping the one a checker judges best, rather than trusting whatever the first attempt produces.

briefingTopics at every level

Writing down everything a model needs in order to act without guessing: the goal, context it lacks, constraints, what a finished result looks like, what to do when unsure, and the output format.

C

calibrating trustalso called trust calibrationTopics at every level

Keeping how much you rely on a model without checking in line with how often it has actually been right on tasks like the one in front of you, tracked per task type, not as one overall impression.

See also: reviewing, eval

capLevel 05 · Agent loops

A hard limit your code enforces on a loop, such as a maximum number of steps or tokens, that forces a stop the model cannot override or even see.

checkpointLevel 03 · Workflows

A saved snapshot of a workflow's shared state, written after a step, so a crashed or interrupted run can resume from that point instead of starting over from the beginning.

chunkLevel 02 · Added context

A passage a document is cut into before indexing, small enough to embed and retrieve on its own, ideally holding one complete idea rather than splitting a sentence or table row in half.

citationLevel 02 · Added context

A pointer from part of an answer back to the specific source passage it came from, so a reader can check whether the source actually supports what the answer claims.

citation hit rateLevel 02 · Added context

The share of questions where every source a grading rule expects was actually cited in the model's answer, used to score how well a retrieval step is working.

code executionLevel 04 · Tool use

Letting the model write a small program instead of choosing among named tools, then running that program in a sandbox your code controls rather than trusting or interpreting it directly.

coding agentLevel 05 · Agent loops

A single agent whose tools read files, edit them, run commands and run tests, looping on a propose-edit-run-test cycle until its own tests pass or a cap ends the run.

See also: sandbox, cap, agent loop

compactionLevel 07 · Always-on agents

Replacing an aging context window with a short written summary once it nears its limit, so a session can carry forward what mattered without keeping the full history.

computer useLevel 04 · Tool use

Letting the model operate a real screen: it looks at a screenshot, picks one action such as a click or a keystroke, your code carries it out, and a new screenshot goes back.

content credentialsalso called C2PALevel 01 · Direct prompting

A record of where a piece of media came from, carried with the file as signed assertions, specified by C2PA rather than by one maker. It says whether that history validates and is free from tampering, not whether the history is good or bad.

context engineeringLevel 02 · Added context

Deciding what goes into a model's request (which instructions, examples, documents and history) and in what order, since the model only knows what it was trained on and what the request contains.

See also: context window, token, prompt caching

context windowLevel 02 · Added context

The amount of text a model can read in one request; material that does not comfortably fit has to be trimmed, retrieved, or summarized before the model ever sees the question.

cosine similarityLevel 02 · Added context

How closely two vectors point in the same direction, used to measure how related two pieces of text are once both have been turned into embeddings.

credential vaultLevel 07 · Always-on agents

Secure storage that holds a password, key or payment method on an agent's behalf, so the agent can use it without ever seeing the raw value itself.

D

delegatingTopics at every level

Deciding which parts of a task to hand to a model and which to keep, based on what a wrong answer would cost, how checkable the result is, and how reversible the action is.

See also: human approval, approval policy, least privilege

distillationTopics at every level

Training a smaller model to imitate a larger model's outputs on a given task, so the smaller one can stand in for the larger one on that same narrow job.

E

embeddingLevel 02 · Added context

A list of floating-point numbers standing in for a piece of text's meaning, positioned so texts with similar meaning get vectors that point in similar directions.

See also: vector, cosine similarity, index

evalalso called evaluationTopics at every level

Running the same fixed set of questions against a system before and after a change, graded the same way both times, so a claim that the change helped can be checked instead of assumed.

See also: golden set, grader, rubric

evaluation frameworkTopics at every level

A tool that holds a dataset of examples, runs a program or a prompt against every one of them, grades each result and lets two runs be compared, instead of a team building that machinery from scratch. Some also record a trace of what happened inside each run.

See also: eval, golden set, grader

F

few-shot examplealso called few-shot learningLevel 01 · Direct prompting

One or more worked examples included in a prompt to show the model a format or pattern rather than only describing it in words.

fine-tuningTopics at every level

Training a model further on your own examples so its behavior on that kind of task becomes more consistent, without repeating the same instructions in every request.

function callingalso called tool call, tool useLevel 04 · Tool use

Giving the model a fixed list of actions your code defined, each with a name, a description and an argument schema, and letting it choose whether to use one, which one, and what arguments to send.

See also: schema, sandbox

G

golden setTopics at every level

A fixed list of questions with a known right answer, or a rubric for judging one, run the same way before and after a change so two runs can be fairly compared.

graderTopics at every level

The thing that scores an answer against a golden set, either by matching it exactly against a pattern or by having another model read it against a rubric, which is itself a judgment call worth checking by hand.

See also: rubric, golden set

graph engineeringThread

One phrase for two different techniques that happen to share a data structure. A knowledge graph connects information: entities and the relationships between them. A workflow graph or an agent graph connects work: steps, and who or what picks the next one.

See also: knowledge graph, agent graph, checkpoint

groundingalso called groundedLevel 02 · Added context

Tying a claim in an answer to a specific source that can actually be checked, the way RAG grounds an answer in retrieved documents instead of whatever the model remembers from training.

See also: citation, RAG

guardrailTopics at every level

A check on what goes into a model or what comes out: an input filter, an output validator, a separate classifier trained to judge safety, a schema check. It is probabilistic and can be wrong in both directions, so it is never the control that holds; a code check that tests a specific fact is.

See also: prompt injection, sandbox, red teaming

H

hallucinationTopics at every level

A fluent, confident answer that is not actually true or not supported by any real source; research argues this happens because training and grading reward a plausible guess over admitting uncertainty.

See also: grounding, reviewing

handoffLevel 06 · Teams of Agents

The edge in an agent graph where one agent's output hands control to another named agent, chosen by a model call rather than a rule your code wrote in advance.

harnessalso called agent harnessLevel 05 · Agent loops

Everything around the model in an agent: the loop that calls it, the tool definitions it is shown and the code that runs them, what goes into its next request, whether an action needs approval, the caps on steps and tokens, the sandbox, and what gets logged. None of it is the model.

See also: agent loop, cap, stop condition, loop engineering

human approvalalso called human-in-the-loopLevel 03 · Workflows

A pause your code inserts before something costly, irreversible or too uncertain to ship, handing the decision to a person instead of letting the run continue on its own.

hybrid searchLevel 02 · Added context

Running a keyword search and a meaning-based search over the same documents and merging the two result lists, so an exact identifier and a paraphrased question can both be found.

I

indexalso called vector indexLevel 02 · Added context

The set of embeddings, or the keyword structure, built from a document set in advance so a later question can be compared against it and ranked, without re-reading every document.

K

knowledge cutoffLevel 01 · Direct prompting

The date after which a model's training data stops, so anything that changed after that date is not something the model actually knows, however confidently it answers.

knowledge graphLevel 02 · Added context

Facts stored as entities and the relationships between them, so a chain of hops can join facts across documents instead of needing one passage to state the whole answer.

See also: retrieval, RAG

L

latencyTopics at every level

How long a request takes to complete, tracked separately from cost; caching, batching and routing to a smaller model are among the main ways to bring it down at volume.

least privilegeTopics at every level

Giving a system only the access its task actually needs, so an action it was never granted is one no instruction, however cleverly worded, can talk it into taking.

levelLevel 05 · Agent loops

One of the eight levels this site sorts a technique into. A new level starts where the answer to “who decides the next step” changes: nobody, you, your code, the model for one action, the model for every step, several models, the models including when to start.

local modelTopics at every level

A model run on your own hardware instead of called over an API, trading a per-call bill for hardware you own and buy, and for models small enough to fit on it.

long-running taskalso called long-horizon taskLevel 07 · Always-on agents

Work that starts on a schedule or an event and continues across many separate sessions until its queue or goal is finished or a person steps in, with no single session seeing the one before it directly.

See also: queue, checkpoint, compaction

loop engineeringLevel 05 · Agent loops

Designing an agent's loop on purpose (what starts it, what it repeats, what stops it) rather than only writing its prompt. Anthropic's Claude Code team defines loops as agents repeating cycles of work until a stop condition is met.

See also: harness, agent loop, stop condition

LoRAalso called low-rank adaptation, adapterTopics at every level

A small, trainable add-on layered onto a model's frozen weights instead of retraining the whole model, cutting the parameters and memory a fine-tuning run needs by orders of magnitude.

M

MCPalso called Model Context ProtocolLevel 04 · Tool use

A standard way for an application to connect to servers that expose tools, resources and prompts to a model, instead of a developer wiring each integration by hand.

See also: function calling, sandbox

memoryLevel 02 · Added context

Keeping information from one conversation to the next by writing facts down somewhere and reading them back into a later, otherwise unrelated conversation, since a chat has no memory of its own.

memory layerLevel 02 · Added context

The part of a product that writes facts down after one conversation and reads them back into a later one. Two different things get called this: a summary, cheap to reread but lossy, and a record kept on its own and searched on demand.

See also: memory, recall, context window

model-decided stepLevel 05 · Agent loops

A step in a recorded run where the model's own output chose what happened next (which tool, which query, whether to stop), as opposed to a step your code decided regardless of what the model said.

See also: trace, cap

multi-agent systemLevel 06 · Teams of Agents

Several agents working on a task, using the same underlying model or different models: a lead agent splitting the work among others, agents handing work to each other across a graph, or two agents checking each other's output. That is level 6 on this site.

See also: orchestrator, agent graph, review and debate

multimodalLevel 01 · Direct prompting

A model that takes or produces more than text (images, audio, video and documents) inside the same one-call request and response shape as an ordinary chat message.

O

observabilityTopics at every level

Recording what each run did in enough detail that a bad result can be traced back to the step that caused it: which passages a retrieval step picked, which tool the model called and with what arguments, which branch a workflow took, and what each step spent.

See also: trace, latency, audit trail

orchestratoralso called lead agentLevel 06 · Teams of Agents

A lead model that reads a task, decides how to split it, hands each piece to a worker, and combines what comes back, rather than doing the whole task itself.

organization of agentsalso called agent swarmLevel 07 · Always-on agents

Several standing agents with distinct roles and a shared goal, where the roster itself keeps running and changing over time rather than being assembled fresh for one job and torn down after.

P

parallel callsalso called parallelization, sectioning, votingLevel 03 · Workflows

Running more than one model call at the same time instead of one after another, then combining the results in code, either splitting one task into independent parts or running the same task several times to vote.

progressive disclosureLevel 05 · Agent loops

Loading only a skill's short name and description into context up front, and its full instructions only once the model actually decides to use it, so unused skills cost almost nothing.

promptLevel 01 · Direct prompting

The request text sent to a model in one call: instructions, examples and the question, everything the model sees that is not already baked into its training.

See also: system prompt, few-shot example

prompt cachingLevel 02 · Added context

Reusing a matching, unchanged prefix of a request across calls at a reduced billing rate, which is why makers recommend putting content that never changes first and content that changes every call last.

prompt chainingLevel 03 · Workflows

Splitting one task into a fixed sequence of steps and handing each step's output to the next, with your code deciding how many steps there are and what gate sits between them.

prompt engineeringLevel 01 · Direct prompting

Writing the request itself well (instructions, examples, an assigned role, a required format, sometimes asking the model to reason first) to get a more consistent result from a single call.

prompt injectionTopics at every level

Text written to look like an instruction, hidden in the user's message or in retrieved or tool-returned content, that tries to redirect what the model does instead of answering the actual question.

See also: guardrail, sandbox

Q

quantizationTopics at every level

Storing a model's weights at lower precision so a large model fits smaller hardware and may run faster. llama.cpp's own documentation says it shrinks the model and can speed up inference, and that it may cost some accuracy. Name the exact level you ran, not just "4-bit".

See also: local model, eval

queueLevel 07 · Always-on agents

A persisted list of pending work items a long-running task drains one at a time across separate sessions, surviving a crash because the queue is checkpointed after every item, not held only in memory.

R

RAGalso called retrieval-augmented generationLevel 02 · Added context

Searching your own documents for the passages closest to a question, putting those passages in the prompt, and asking the model to answer once using only what it was given.

See also: retrieval, embedding, chunk, agentic RAG, citation

RAG 2.0Level 02 · Added context

A label rather than a defined technique: different people use it for different bundles of retrieval improvements: better chunking, reranking, hybrid search, retrieval the model itself drives. This site has no page under that name; the pages on RAG, embeddings and search, and agentic RAG cover the substance.

See also: RAG, reranking, hybrid search, agentic RAG

reasoning effortalso called extended thinkingLevel 01 · Direct prompting

How much a model is allowed to work through a problem in its own words before answering; makers recommend more of it for math, debugging and planning and less for simple lookups.

reasoning modelalso called thinking modelLevel 01 · Direct prompting

A model trained, usually with reinforcement learning, to write out a long chain of thought before it answers, and to notice and fix its own mistakes along the way. OpenAI's o1 (September 2024) was the first sold as one.

See also: reasoning effort, test-time compute

recallLevel 02 · Added context

The operation that ranks stored facts or documents against a new question and returns the closest matches, the same mechanism whether the store is a document corpus or a small set of personal facts.

red teamingTopics at every level

Attacking your own system on purpose, before someone else does it without permission, across the system around the model and not only the model itself. What matters is not the report but the change it forces, most durably a test that keeps failing until the finding is fixed.

See also: guardrail, prompt injection, eval

rerankingLevel 02 · Added context

A second pass over a first, larger cut of retrieved candidates, scored by a model trained specifically to sort results by relevance to the original query.

retrievalLevel 02 · Added context

Searching an index of document chunks for the ones closest to a question, and keeping a fixed number of the best matches to hand to the model.

review and debatealso called debate and review, multi-agent debateLevel 06 · Teams of Agents

A separate agent, with its own context and often its own retrieval, checks or argues with another agent's work and decides whether to accept it, instead of one fixed, code-owned test.

See also: write and check

reviewingTopics at every level

Checking work you did not do yourself before it goes anywhere: pulling out specific claims and checking them against their source, then checking what is missing, then judging the result.

See also: hallucination, citation

routingLevel 03 · Workflows

Looking at an input, deciding which of several fixed kinds it is, and sending it down the handler built for that kind, with a real fallback for whatever fits none of them.

See also: parallel calls, structured output

rubricTopics at every level

A written checklist a grader model reads an answer against when the answer is too open-ended to match against a fixed pattern, used to turn a judgment call into a repeatable score.

S

sandboxLevel 04 · Tool use

An isolated environment with no network access and fixed resource limits, where code the model wrote is actually run, so what the model produces is data your code hands to an interpreter, never code it trusts directly.

schemaLevel 01 · Direct prompting

A fixed shape for a reply (named fields with defined types) that a model's output is constrained to match, so downstream code can parse it without guessing at its structure.

self-consistencyLevel 01 · Direct prompting

Asking the same question several times as independent calls and returning whichever answer the largest share of the samples agree on, instead of trusting a single attempt.

skillLevel 05 · Agent loops

Instructions an agent keeps on the shelf until it decides it needs them, loaded into context only when triggered, instead of text repeated into every single turn like a system prompt.

See also: progressive disclosure, system prompt

stateless MCPalso called sessionless MCPLevel 04 · Tool use

The current MCP specification, revision 2026-07-28, removed the initialize handshake and protocol-level sessions. Every request now carries its own protocol version and capabilities, and a server that needs state across calls returns an explicit, server-minted handle the client passes back as an ordinary tool argument.

See also: MCP, function calling

stop conditionLevel 05 · Agent loops

The test a model-driven loop uses to decide it is finished and should give a final answer rather than take another action; a cap can force a stop the model never actually reaches.

structured outputLevel 01 · Direct prompting

A reply constrained to come back in a fixed shape, such as a JSON object with named fields, instead of a paragraph your code has to parse by guessing.

synthetic dataTopics at every level

Training examples generated with a model rather than collected from real use, for a later fine-tuning or evaluation run.

System One modelLevel 01 · Direct prompting

TypeSafe AI's name for a model that generates no text and returns a typed decision with a probability in one fast pass, for use inside software. Jev (September 2026) is the first. The name borrows psychology's fast System 1, as against slow, deliberate System 2.

See also: structured output, reasoning model

system promptLevel 01 · Direct prompting

The instructions sent to a model at the start of a request, separate from what the user or the retrieved content says, setting how it should behave for that call.

T

test-time computealso called inference-time computeLevel 01 · Direct prompting

Computation spent while answering, as opposed to while training. Reasoning models made it a dial: more thinking time, better answers on hard problems, at the cost of seconds and output tokens.

See also: reasoning model, reasoning effort

tokenLevel 02 · Added context

The unit a model's input and output are measured and billed in; every cost strip on this site counts tokens in and tokens out for the run it illustrates.

traceLevel 05 · Agent loops

A recorded, step-by-step account of a run (every model call, every tool call, and whether each step was decided by code or by the model) used to explain how an answer was actually produced.

See also: model-decided step, eval

trajectoryTopics at every level

The path a run took to reach its answer: which tools were called, in what order, with what arguments. Scoring it is a separate measurement from scoring the final answer, and it gives partial credit for the steps a run got right.

See also: trace, eval, agent loop

V

vectorLevel 02 · Added context

A fixed-length list of numbers representing a piece of text, an embedding, positioned in space so that texts with related meaning end up close together.

vector databasealso called vector storeLevel 02 · Added context

A database built to hold embeddings and rank them by similarity to a query vector: pgvector, Pinecone, Weaviate and Qdrant are four. It stores the index a retrieval step searches; the embeddings it stores come from an embedding model, not from the database.

See also: embedding, index, vector

vision-language-action modelalso called VLALevel 07 · Always-on agents

A model that converts vision and language input directly into motor control, letting a robot take a physical action rather than only produce text.

voice agentLevel 05 · Agent loops

A single agent you talk to instead of type to, in real time, that additionally has to decide when a caller has stopped talking and what to do if they interrupt.

W

write and checkalso called evaluator-optimizerLevel 03 · Workflows

A loop where one prompt writes a draft and a separate prompt checks it against one specific, testable criterion, revising and rechecking until it passes or a fixed revision cap is reached.

See also: review and debate, rubric