Beyond the labels

The distinctions that change the design.

The levels are this guide’s teaching framework, not an industry standard or a mandatory ladder. Real systems combine patterns. Choose the capability that solves a demonstrated problem.

01

Context, retrieval, and memory

What information must be available for this request?

Send a small, relevant source packet. Add retrieval when the collection is too large or changes too often to include directly.

A long context window is capacity, not a relevance or accuracy guarantee. Retrieval selects information; memory persists selected information across requests. A system may use all three.
Tradeoffs and what to measure

Preserve document identity, access scope, applicability dates, and source provenance. Decide what memory may be written, corrected, expired, or deleted. A summary can lose a constraint; a stale memory can contradict a current source. Test both missing evidence and misleading evidence.

Check it: Measure retrieval coverage separately from whether the answer is supported. Include questions with no answer and contradictory sources.

02

Workflow or agent?

Can you specify the useful steps before the run starts?

Use a fixed workflow when the sequence and branch rules are known. Use an agent loop when the next useful action depends on what the model discovers.

The presence of model calls, conditional branches, or a framework named “agent” does not settle this. Look at who selects the next action in the actual system.
Tradeoffs and what to measure

Real applications are often hybrids: a deterministic intake flow can invoke a bounded research agent, validate its output, and wait for approval. Keep controls in code even when task planning is model-directed.

Check it: Compare completion quality, failure rate, latency, and cost against the simpler workflow on the same tasks.

03

Reasoning effort or more tools?

Is the bottleneck deciding well, or missing an observation?

Give the model adequate evidence and useful tools, then tune reasoning effort against your actual tasks.

Reasoning is important for planning, diagnosing failures, and revising a course of action. A dedicated reasoning model can improve those decisions, but “reasoning model” and “agent” describe different things.
Tradeoffs and what to measure

More inference-time compute can increase latency and cost without resolving absent or wrong evidence. A request can reason extensively and never act; an agent can act through a tool loop with a model that exposes no separate reasoning setting. Inspect outcomes and observable actions rather than demanding private chain-of-thought.

Check it: Compare successful outcomes and tool choices across available reasoning settings. Record cost and time; do not use answer length as a proxy for reasoning quality.

04

One agent or several?

Does decomposition reduce a real bottleneck?

Try one capable, well-equipped agent first. Split work when subtasks can usefully proceed independently or need separate context and permissions.

Parallel execution of a fixed plan is different from a lead agent deciding what to delegate. Separate agent instances do not guarantee independent mistakes.
Tradeoffs and what to measure

Define task ownership, a handoff contract, the evidence a worker must return, and what happens on timeout or disagreement. Avoid letting every worker modify the same shared artifact. A reviewer using the same misleading sources can reinforce an error instead of catching it.

Check it: Include coordination overhead, duplicated searches, omitted findings, and synthesis errors in the comparison—not just wall-clock speed.

05

Structured output or verified truth?

What does the downstream code actually need to trust?

Define the output contract, validate it, and recompute things ordinary code can check.

Valid JSON, schema compliance, factual accuracy, and permission to act are four separate properties. Passing one does not establish the others.
Tradeoffs and what to measure

Plan for missing values, refusals, truncated responses, extra fields, and invalid tool arguments. A citation validator can establish that a source ID exists, while a separate check asks whether the source supports the claim. Automatic repairs must not invent missing facts.

Check it: Track format failures and factual errors separately. Keep a source-backed reference or a checkable external outcome for important fields.

06

Tools, protocols, and authority

What is the system allowed to do, and who enforces that?

Expose the smallest useful tool set and validate arguments and authorization before execution.

MCP standardizes a connection to tools and context. It does not turn a tool description into authorization, validate every business rule, or make retrieved instructions trustworthy.
Tradeoffs and what to measure

Keep identity, scope, consent, and execution checks in the host application. Treat tool results and retrieved documents as potentially untrusted. A prompt-injection defense is stronger when a malicious instruction cannot acquire a capability the tool layer never granted.

Check it: Exercise denied tools, out-of-scope arguments, untrusted instructions, changed approvals, and uncertain external outcomes.

07

Long-running or always-on?

Must work survive a session ending, a restart, or a quiet period?

Save a durable checkpoint and define how a new run starts, resumes, stops, and asks for help.

Long-horizon describes work spanning time. Always-on describes availability and triggers. Neither requires a team of agents or constant model generation.
Tradeoffs and what to measure

A context summary is not the same as a durable task ledger or an external-action receipt. Preserve completed work, remaining work, blockers, relevant evidence, and a reproducible environment. A crash between an external effect and recording its receipt creates an ambiguous retry; resolve that ambiguity before repeating the action.

Check it: Test process interruption, stale checkpoints, replayed events, expired authority, and duplicate actions. Verify the actual artifact after resuming.

08

Evaluation, observation, and release

What evidence would justify relying on this system?

Define a representative test set and acceptance criteria, then inspect failures alongside aggregate scores.

An evaluation grades task behavior or outcomes. A trace records how a run unfolded. Both are useful; neither replaces the other.
Tradeoffs and what to measure

Include typical cases, difficult cases, and cases where the correct behavior is to refuse, abstain, or ask for help. Keep prompt-development cases separate from held-out evaluation. Repeat stochastic tasks, calibrate model graders against human judgment, and check final environment state rather than trusting a success message.

Check it: Report task success, unsafe or unsupported outcomes, latency, resource use, and uncertainty. Record the model, harness, tools, data, and configuration so comparisons are meaningful.

Try the idea with your AI.

The practical recipe examples include a copyable brief, sample inputs, and review criteria. Start with a small case whose correct behavior you can explain before adapting it to your own material.

Try a recipe →