You supply the request and decide what to do with the reply.
A focused everyday life example. Additional perspectives appear where they provide a useful contrast.
GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions
Chat: see it in practice.
A conversational interface in which a person supplies messages and decides what to do with the replies.
What you’ll walk through
Follow a request through a first response and a correction. Watch how a useful conversation separates supplied facts, proposed wording, and details that still need an answer.
The task in this version
Draft an invitation for our repair workshop. Ask before filling in missing facts.
What you’ll learn to check
A checked facts list and a revised draft grounded in information the user actually provided.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Everyday lifeAn authored case with its own evidence, changed condition, and decision.
The task in this example
Draft an invitation for our repair workshop. Ask before filling in missing facts.
Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Known: Saturday, free admission, bring one broken item. Venue: not supplied.
What changed: Establish the facts supplied for this version of the task.
WHY THIS MATTERS
What this case assumes
A conversational reply may sound confident even when the request is incomplete. Decide which gaps affect correctness and which can remain editable suggestions.
1 / 6
Apply this to your project
Describe your task to your own model and use Chat as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
Go deeper: practical guidance, failure modes, and implementation
This page’s level 1 example is a plain model call: one request and one response, with no search
or tool execution around it. You supply a message and inspect the reply. A chat interface does
not guarantee that architecture: a modern chat product may search, run tools, use memory, or
coordinate several model calls behind one visible reply. The level describes how the task runs,
not the appearance of its message box.
Because there is only one step, the whole outcome depends on what happens inside that one call:
how the model was trained, and what you put in the request. Later levels add retrieval, tools and
loops around the model to make up for what one call alone gets wrong. Before any of that, it
helps to be precise about what “the model” even refers to, because a chat app, the company behind
it, and the model actually answering you are three different things often sharing one name.
This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.
Optional: inspect the implementation trace
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
Chat: one call, no documents
Send the question straight to the model, with no retrieval and no tools, and return whatever comes back.
Level 1 · Direct prompting
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step
The run, step by step
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
STEP 01 / 03Your code chose
The question arrives
"What voltage does a DR-210 need?"
0 tokens · 0 ms
Practical guidance
The name that actually matters is not “Claude” or “ChatGPT,” it’s the specific model your product
used to answer you. Most chat apps show this in a menu near the message box, a dropdown at the top
of the conversation, or a settings panel; look there before anywhere else. Anthropic’s own
documentation names Claude Sonnet 5 as a specific, versioned model, released June 30, 2026 with the
model id claude-sonnet-5[1]: that is the kind of name the menu is naming, not the
product’s own name.
When a colleague’s advice about “Claude” doesn’t reproduce, the product name isn’t enough
information to debug from. Ask which model their menu showed, not which app they opened: Claude,
Anthropic’s own product, is described as “a helpful, intuitive, and powerful collaborator you can
put to work on real tasks”[2], and that description covers more than one model
underneath it. ChatGPT is the same shape: one product name over several possible models. If
neither of you checked the menu, you were never comparing the same thing in the first place.
Two things to check directly in the product itself, not by asking a model about itself from
memory. Type “What model are you, and what is your knowledge cutoff?” and compare the answer
against the menu’s own label and the product’s own documentation; a model can be wrong about its
own name and date, and the menu is the source that actually decides. And before comparing “Claude”
against another product’s newest model by name, check that you are comparing two models rather
than a product against a model, since a product can run more than one and the menu decides which.
The developer and the tool matter mainly when you’re building something, not chatting. OpenAI’s
own documentation describes its API as what lets a developer “prompt a model and generate text” or
“build agents that use tools and computers”[3]: if a feature was clearly built by
someone else’s code calling a model, that’s the tool layer, and not something you troubleshoot
from inside a chat window. Ollama states plainly that “Nothing you run locally ever leaves your
machine”[4], which matters only if privacy is the actual question you came with.
None of this matters for a one-off question you can check yourself. It matters when a result
doesn’t match what you expected, or what someone else got.
Implementation details
The Build it example is level 1’s entire trick: send the question, change nothing else, and see
what the model does with no documents and no tools available to it. examples/one_call asks
about Halvorsen, a fictional appliance maker invented for this site’s synthetic documents,
specifically so a model has never legitimately seen its manuals. A correct run at this level
mostly means declining questions it cannot know the answer to, rather than inventing a
plausible-sounding number. That is exactly the failure this level exists to measure, and
exactly what RAG exists to fix.
The system prompt is the only lever available here, and it says so directly: answer plainly, and
say so plainly when a specific fact is not known instead of guessing at it. There is no chunking,
no search and no schema: one system message, one user message, one call.
Two things are worth noticing in the trace. First, calling the model is not itself a model
decision: the code decided to make this call, in this order, before the model said anything: the
"model" kind on that step and the "code"decided_by on the same step answer two different
questions (see docs/EVALS.md). Second, there is nothing left to decide once the call returns;
the code does not parse the answer, check it against anything, or call the model again. It hands
back exactly what came back. Run it yourself:
examples/one_call/README.md · lines 15–15
python -m examples.one_call --model stub:scripted
A newcomer building their first real thing on top of a model almost always starts here, whether or
not they call it “level 1”: one prompt, sent through whichever tool wraps the developer’s model,
usually that developer’s own API, or a tool like Ollama that can swap which model answers without
changing the calling code.
When you do not need this
Try level 0, no model at all first if the question has a fixed
vocabulary and repeats often enough that a keyword search or a rule can answer it with no model at
all. Level 0 is fast, free and completely predictable: three things a chat reply cannot
promise.
Completion in the editor
The editor that finishes your line as you type is this level too, and for people who write code it
is usually the model they touch most hours of the week. GitHub’s documentation says “GitHub
Copilot offers coding suggestions as you type”[5]. Cursor describes its own the same
way: “Tab is Cursor’s AI-powered autocomplete. It suggests code as you type, based on your recent
edits, surrounding code, and linter errors”[6].
Nobody decides the next step, which is what keeps it on this rung. The editor’s code decides when
to ask and what context to send, the model fills in the rest of the line, and you accept it with a
keystroke or keep typing. That is one request and one answer. It is not level 5: a
coding agent picks each step and decides when it is
finished, and a completion picks nothing.
You do not need it when you already know exactly what the line says, since reading a suggestion
costs more attention than typing eight characters, and at a bench a register write that looks
right is worse than a blank line.
The failure mode follows from that asymmetry: accepting takes one key and checking takes a minute,
so a plausible wrong line lands in the file unread. There is a second thing an unread line can
carry: GitHub’s documentation says “GitHub Copilot checks each suggestion for matches with
publicly available code”[7], and that a match is either discarded or offered with a code
reference, depending on a policy setting your account or organization controls[7].
Failure modes
Confident answers outside what the model actually knows
How to notice it
The reply is fluent and specific about something the model was never trained on (a fictional product, your own private data, an internal document), instead of saying it does not know.
How to test for it
Ask about something invented for this site's synthetic corpus, like a Halvorsen part number, with no documents attached, and check whether the model declines or guesses.
No memory beyond what is sent
How to notice it
A follow-up question gets answered as if the earlier part of the conversation never happened, because a single call only sees what is in that one request.
How to test for it
Call the model with only the latest question, no prior turns included, and check whether it can still answer something that depended on earlier context.
Knowledge cutoff
How to notice it
The model answers confidently about something that changed after its training data ends, using the old fact as if it were current.
How to test for it
Ask about a recent event or a fact you know changed recently, and compare the answer against the model's stated knowledge cutoff.
No way to check its own answer
How to notice it
Asking "are you sure" is still just another single call; the model may double down or flip its answer with equal confidence either way, since nothing verifies either reply against a source.
How to test for it
Ask the same factual question twice in separate calls, phrased differently, and check whether the two answers actually agree.
Cost and latency
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
1Model calls, one question
~40Tokens in
~55Tokens out
~0.6sWall time
Compared with RAG (level 2)RAG adds one retrieval step and roughly forty times the input tokens for the same question, in exchange for grounding the answer in real documents instead of whatever the model remembers from training.
Chat is scored on the same 60-question set as every other level (docs/EVALS.md). With no
documents attached, its lookup and numeric scores should sit close to level 0’s floor: whatever
it gets right, it gets right from training data alone, which for a fictional appliance maker like
Halvorsen should be close to nothing. The one place a single call can beat a keyword score is
unanswerable questions, if the system prompt’s instruction to decline rather than guess actually
holds: it can say “I don’t know” in its own words instead of returning an irrelevant passage.
No result file exists yet for any level (see docs/EVALS.md). Run
python scripts/eval_run.py --example one_call --model stub --dry to project the token cost of a
run before spending anything on a real one.
Run it
What to monitor
The rate of confidently wrong answers on anything outside common knowledge, since a single call has no way to flag its own uncertainty beyond what the prompt asks it to say.
Cost at volume
Tokens in track what you send (the question plus any instructions); tokens out track how long the replies run. Both scale linearly with traffic, and there is no retrieval or tool infrastructure running alongside it to add to the bill.
How it fails in production
A user asks about something the model was never trained on, or something that changed after its training cutoff, and gets a fluent, wrong answer instead of a refusal. Or an app update quietly drops the instruction to say "I don't know", and nobody notices until a wrong answer causes a real problem.
What to log
The full prompt sent (system and user messages), the model id and version, and the raw reply, so a bad answer traces back to what the model was actually asked rather than being guessed at afterward.
Try it
Use it
Open a chat app and check which model answered your last message, usually in a menu or settings panel. Ask it its own knowledge cutoff date and compare that against the product's documentation.
Build it
Run python -m examples.one_call --model stub:scripted from the repo root. With no documents and no tools, the reply declines to give the DR-210's supply voltage and says where the number is: the best answer this level has. Run it again with --model stub: the echo prints the prompt back, the same shape, nothing in it.
Either lane
Pick a name from the "Out there" list at the bottom of this page: which of the four kinds is it, developer, model, product or tool?
Build it
At a bench, paste a paragraph from an instrument programming manual into a chat app and ask for a summary: a safe use, since the text is right there to check. Then ask it, with nothing attached, for a specific accuracy figure from memory. A fluent answer is not a reported measurement, and never becomes one.