Level 01 · Direct prompting

Chat

Asking a model a question in a chat app.

Sourced

Concept at a glance

One message in. One reply out.

SequenceConceptual illustration
One message in. One reply out.Your message leads to Model. Model leads to Your next move. You supply the request and decide what to do with the reply.Your messageQuestion and contextModelGenerates a responseYour next moveUse, edit, or ask againOne message in. One reply out.Your message leads to Model. Model leads to Your next move. You supply the request and decide what to do with the reply.Your messageQuestion and contextModelGenerates a responseYour next moveUse, edit, or ask again
Read the connections in words
  • Your message → Model: Generates a response.
  • Model → Your next move: Use, edit, or ask again.
Key idea

You supply the request and decide what to do with the reply.

A focused everyday life example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Chat: see it in practice.

A conversational interface in which a person supplies messages and decides what to do with the replies.

What you’ll walk through

Follow a request through a first response and a correction. Watch how a useful conversation separates supplied facts, proposed wording, and details that still need an answer.

The task in this version

Draft an invitation for our repair workshop. Ask before filling in missing facts.

What you’ll learn to check

A checked facts list and a revised draft grounded in information the user actually provided.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Everyday lifeAn authored case with its own evidence, changed condition, and decision.
The task in this example

Draft an invitation for our repair workshop. Ask before filling in missing facts.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Known: Saturday, free admission, bring one broken item. Venue: not supplied.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

A conversational reply may sound confident even when the request is incomplete. Decide which gaps affect correctness and which can remain editable suggestions.

1 / 6

Apply this to your project

Describe your task to your own model and use Chat as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

This page’s level 1 example is a plain model call: one request and one response, with no search or tool execution around it. You supply a message and inspect the reply. A chat interface does not guarantee that architecture: a modern chat product may search, run tools, use memory, or coordinate several model calls behind one visible reply. The level describes how the task runs, not the appearance of its message box.

Because there is only one step, the whole outcome depends on what happens inside that one call: how the model was trained, and what you put in the request. Later levels add retrieval, tools and loops around the model to make up for what one call alone gets wrong. Before any of that, it helps to be precise about what “the model” even refers to, because a chat app, the company behind it, and the model actually answering you are three different things often sharing one name.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Chat: one call, no documents

Send the question straight to the model, with no retrieval and no tools, and return whatever comes back.

Level 1 · Direct prompting
QuestionQuestionBuild the promptBuild the promptMODELanswers onceanswers onceAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 03Your code chose

The question arrives

"What voltage does a DR-210 need?"
0 tokens · 0 ms

Practical guidance

The name that actually matters is not “Claude” or “ChatGPT,” it’s the specific model your product used to answer you. Most chat apps show this in a menu near the message box, a dropdown at the top of the conversation, or a settings panel; look there before anywhere else. Anthropic’s own documentation names Claude Sonnet 5 as a specific, versioned model, released June 30, 2026 with the model id claude-sonnet-5[1]: that is the kind of name the menu is naming, not the product’s own name.

When a colleague’s advice about “Claude” doesn’t reproduce, the product name isn’t enough information to debug from. Ask which model their menu showed, not which app they opened: Claude, Anthropic’s own product, is described as “a helpful, intuitive, and powerful collaborator you can put to work on real tasks”[2], and that description covers more than one model underneath it. ChatGPT is the same shape: one product name over several possible models. If neither of you checked the menu, you were never comparing the same thing in the first place.

Two things to check directly in the product itself, not by asking a model about itself from memory. Type “What model are you, and what is your knowledge cutoff?” and compare the answer against the menu’s own label and the product’s own documentation; a model can be wrong about its own name and date, and the menu is the source that actually decides. And before comparing “Claude” against another product’s newest model by name, check that you are comparing two models rather than a product against a model, since a product can run more than one and the menu decides which.

The developer and the tool matter mainly when you’re building something, not chatting. OpenAI’s own documentation describes its API as what lets a developer “prompt a model and generate text” or “build agents that use tools and computers”[3]: if a feature was clearly built by someone else’s code calling a model, that’s the tool layer, and not something you troubleshoot from inside a chat window. Ollama states plainly that “Nothing you run locally ever leaves your machine”[4], which matters only if privacy is the actual question you came with.

None of this matters for a one-off question you can check yourself. It matters when a result doesn’t match what you expected, or what someone else got.

Implementation details

The Build it example is level 1’s entire trick: send the question, change nothing else, and see what the model does with no documents and no tools available to it. examples/one_call asks about Halvorsen, a fictional appliance maker invented for this site’s synthetic documents, specifically so a model has never legitimately seen its manuals. A correct run at this level mostly means declining questions it cannot know the answer to, rather than inventing a plausible-sounding number. That is exactly the failure this level exists to measure, and exactly what RAG exists to fix.

The system prompt is the only lever available here, and it says so directly: answer plainly, and say so plainly when a specific fact is not known instead of guessing at it. There is no chunking, no search and no schema: one system message, one user message, one call.

examples/one_call/run.py · lines 22–39
def run(question: str, model: Model, embedder: Embedder | None, tracer: Tracer) -> Answer:
    del embedder  # level 1 has no retrieval step
    messages = [
        Message(role="system", content=SYSTEM_PROMPT),
        Message(role="user", content=question),
    ]
    tracer.record(kind="code", decided_by="code", title="Build prompt", detail=question)
    completion = model.complete(messages, max_tokens=400)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask the model",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    return Answer.from_text(completion.text)

Two things are worth noticing in the trace. First, calling the model is not itself a model decision: the code decided to make this call, in this order, before the model said anything: the "model" kind on that step and the "code" decided_by on the same step answer two different questions (see docs/EVALS.md). Second, there is nothing left to decide once the call returns; the code does not parse the answer, check it against anything, or call the model again. It hands back exactly what came back. Run it yourself:

examples/one_call/README.md · lines 15–15
python -m examples.one_call --model stub:scripted

A newcomer building their first real thing on top of a model almost always starts here, whether or not they call it “level 1”: one prompt, sent through whichever tool wraps the developer’s model, usually that developer’s own API, or a tool like Ollama that can swap which model answers without changing the calling code.

When you do not need this

Try level 0, no model at all first if the question has a fixed vocabulary and repeats often enough that a keyword search or a rule can answer it with no model at all. Level 0 is fast, free and completely predictable: three things a chat reply cannot promise.

Completion in the editor

The editor that finishes your line as you type is this level too, and for people who write code it is usually the model they touch most hours of the week. GitHub’s documentation says “GitHub Copilot offers coding suggestions as you type”[5]. Cursor describes its own the same way: “Tab is Cursor’s AI-powered autocomplete. It suggests code as you type, based on your recent edits, surrounding code, and linter errors”[6].

Nobody decides the next step, which is what keeps it on this rung. The editor’s code decides when to ask and what context to send, the model fills in the rest of the line, and you accept it with a keystroke or keep typing. That is one request and one answer. It is not level 5: a coding agent picks each step and decides when it is finished, and a completion picks nothing.

You do not need it when you already know exactly what the line says, since reading a suggestion costs more attention than typing eight characters, and at a bench a register write that looks right is worse than a blank line.

The failure mode follows from that asymmetry: accepting takes one key and checking takes a minute, so a plausible wrong line lands in the file unread. There is a second thing an unread line can carry: GitHub’s documentation says “GitHub Copilot checks each suggestion for matches with publicly available code”[7], and that a match is either discarded or offered with a code reference, depending on a policy setting your account or organization controls[7].

Failure modes

Confident answers outside what the model actually knows

How to notice it
The reply is fluent and specific about something the model was never trained on (a fictional product, your own private data, an internal document), instead of saying it does not know.
How to test for it
Ask about something invented for this site's synthetic corpus, like a Halvorsen part number, with no documents attached, and check whether the model declines or guesses.

No memory beyond what is sent

How to notice it
A follow-up question gets answered as if the earlier part of the conversation never happened, because a single call only sees what is in that one request.
How to test for it
Call the model with only the latest question, no prior turns included, and check whether it can still answer something that depended on earlier context.

Knowledge cutoff

How to notice it
The model answers confidently about something that changed after its training data ends, using the old fact as if it were current.
How to test for it
Ask about a recent event or a fact you know changed recently, and compare the answer against the model's stated knowledge cutoff.

No way to check its own answer

How to notice it
Asking "are you sure" is still just another single call; the model may double down or flip its answer with equal confidence either way, since nothing verifies either reply against a source.
How to test for it
Ask the same factual question twice in separate calls, phrased differently, and check whether the two answers actually agree.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1Model calls, one question
~40Tokens in
~55Tokens out
~0.6sWall time
Compared with RAG (level 2)RAG adds one retrieval step and roughly forty times the input tokens for the same question, in exchange for grounding the answer in real documents instead of whatever the model remembers from training.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

Chat is scored on the same 60-question set as every other level (docs/EVALS.md). With no documents attached, its lookup and numeric scores should sit close to level 0’s floor: whatever it gets right, it gets right from training data alone, which for a fictional appliance maker like Halvorsen should be close to nothing. The one place a single call can beat a keyword score is unanswerable questions, if the system prompt’s instruction to decline rather than guess actually holds: it can say “I don’t know” in its own words instead of returning an irrelevant passage.

No result file exists yet for any level (see docs/EVALS.md). Run python scripts/eval_run.py --example one_call --model stub --dry to project the token cost of a run before spending anything on a real one.

Run it

What to monitor

The rate of confidently wrong answers on anything outside common knowledge, since a single call has no way to flag its own uncertainty beyond what the prompt asks it to say.

Cost at volume

Tokens in track what you send (the question plus any instructions); tokens out track how long the replies run. Both scale linearly with traffic, and there is no retrieval or tool infrastructure running alongside it to add to the bill.

How it fails in production

A user asks about something the model was never trained on, or something that changed after its training cutoff, and gets a fluent, wrong answer instead of a refusal. Or an app update quietly drops the instruction to say "I don't know", and nobody notices until a wrong answer causes a real problem.

What to log

The full prompt sent (system and user messages), the model id and version, and the raw reply, so a bad answer traces back to what the model was actually asked rather than being guessed at afterward.

Try it

  1. Use it

    Open a chat app and check which model answered your last message, usually in a menu or settings panel. Ask it its own knowledge cutoff date and compare that against the product's documentation.

  2. Build it

    Run python -m examples.one_call --model stub:scripted from the repo root. With no documents and no tools, the reply declines to give the DR-210's supply voltage and says where the number is: the best answer this level has. Run it again with --model stub: the echo prints the prompt back, the same shape, nothing in it.

  3. Either lane

    Pick a name from the "Out there" list at the bottom of this page: which of the four kinds is it, developer, model, product or tool?

  4. Build it

    At a bench, paste a paragraph from an instrument programming manual into a chat app and ask for a summary: a safe use, since the text is right there to check. Then ask it, with nothing attached, for a specific accuracy figure from memory. A fluent answer is not a reported measurement, and never becomes one.

How it connects

Before, after and instead of this

Move up when

Optional: products, tools, and models

43 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 37 more examples
In practice

Rewrite a short email

Paste your draft, ask for a clearer version, and review the reply before sending it.

Out there

Named products, tools and models

Products10
  • ChatGPTOpenAI · chat app
  • ClaudeAnthropic · chat app
  • CursorAnysphere · coding agent in an editor
  • DeepSeekDeepSeek · chat app
  • GeminiGoogle · chat app
  • GitHub CopilotGitHub · coding agent in an editor
  • GrokSpaceXAI · chat app
  • Meta AIMeta · chat app
  • Microsoft CopilotMicrosoft · chat app
  • Mistral VibeMistral AI · ai agent for work and coding · formerly Le Chat
Tools10
  • AI SDKVercel · TypeScript AI and agent SDK
  • Claude APIAnthropic · model API
  • Gemini APIGoogle · model API
  • LiteLLMBerriAI · one API for many models
  • llama.cppopen source · runs models locally
  • LM StudioElement Labs · runs models locally
  • OllamaOllama · runs models locally
  • OpenAI APIOpenAI · model API
  • OpenRouterOpenRouter · one API for many models
  • TransformersHugging Face · model library
Models25
  • Claude Haiku 4.5Anthropic · small model
  • Claude Opus 5Anthropic · frontier model
  • Claude Sonnet 5Anthropic · mid-size model
  • Command A+Cohere · enterprise model · formerly Command, deprecated September 15, 2025
  • DeepSeek V4DeepSeek · open-weight modelSuperseded by DeepSeek-V4.1-Flash
  • DeepSeek-V4.1-FlashDeepSeek · open-weight model
  • Gemini 3.1 ProGoogle · frontier model
  • Gemini 3.5 Flash-LiteGoogle · small model
  • Gemini 3.8 FlashGoogle · mid-size model
  • Gemma 4Google · open-weight model · formerly Gemma
  • GLM-5.3Z.ai · open-weight model
  • GPT-5.6 LunaOpenAI · small model
  • GPT-5.6 SolOpenAI · frontier modelSuperseded by GPT-6 Astra
  • GPT-5.6 TerraOpenAI · mid-size model
  • GPT-6 AstraOpenAI · frontier model
  • gpt-ossOpenAI · open-weight model
  • Grok 4.6SpaceXAI · frontier model
  • Kimi K3Moonshot AI · open-weight model
  • Llama 4Meta · open-weight model · formerly Llama
  • Mistral Large 3Mistral AI · open-weight model
  • Mistral Small 4Mistral AI · open-weight model
  • Muse Spark 1.3Meta · frontier model · formerly Muse Spark, 2026-04
  • Nemotron 3NVIDIA · open-weight model · formerly Nemotron, superseded December 15, 2025
  • Phi-4-miniMicrosoft · small open-weight model · formerly Phi
  • Qwen3.8Alibaba · open-weight model

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Claude Sonnet 5 · Anthropic, 06/30/2026 (accessed 09/19/2026)
  2. Claude · Anthropic (accessed 09/19/2026)
  3. OpenAI API Platform Documentation · OpenAI (accessed 09/19/2026)
  4. Ollama · Ollama (accessed 09/19/2026)
  5. Getting code suggestions in your IDE with GitHub Copilot · GitHub (accessed 09/19/2026)
  6. Tab completion · Cursor (accessed 09/19/2026)
  7. Code suggestions · GitHub (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page