# Chat

_Level 01 · Direct prompting · sourced_

Asking a model a question in a chat app.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a request through a first response and a correction. Watch how a useful conversation separates supplied facts, proposed wording, and details that still need an answer.

**Assumptions:** A conversational reply may sound confident even when the request is incomplete. Decide which gaps affect correctness and which can remain editable suggestions.

**Design choices:** Ask about consequential unknowns; make labeled, reversible suggestions for preferences. You do not need an approval workflow for every draft sentence.

**Request:** Draft an invitation for our repair workshop. Ask before filling in missing facts.

**Starting evidence:** Known: Saturday, free admission, bring one broken item. Venue: not supplied.

**Action and control:** Draft from the supplied facts and leave the venue unresolved; the user decides what happens next.

**Stage records (authored, not executed):**

### Input record

Known: Saturday, free admission, bring one broken item. Venue: not supplied.

What changed: Establish the facts supplied for this version of the task.

### Design note

Ask about consequential unknowns; make labeled, reversible suggestions for preferences. You do not need an approval workflow for every draft sentence.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Draft from the supplied facts and leave the venue unresolved; the user decides what happens next.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Join our free repair workshop this Saturday. Bring one broken item. Venue: please confirm before sharing.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

A checked facts list and a revised draft grounded in information the user actually provided.

If the result falls short:
Correct a mistaken assumption explicitly and ask for a revised response. Recheck retained facts rather than assuming the correction fixed everything.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use the pattern for an explanation, plan, invitation, or brainstorming session. Choose what a helpful first draft should accomplish and what you will check yourself.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Join our free repair workshop this Saturday. Bring one broken item. Venue: please confirm before sharing.

**Change something — Supply the venue:** Updated draft: Join us at Oak Hall this Saturday for a free repair workshop. Bring one broken item. This is a draft, not a sent message.

**Decision:** Did the assistant know the venue before you supplied it?

**Answer:** No; it was missing from the context.

**Why:** Contrast a useful draft with an invented venue detail; the conversation has no implicit access to your calendar or facts.

**Review criteria:** A checked facts list and a revised draft grounded in information the user actually provided.

**Recovery:** Correct a mistaken assumption explicitly and ask for a revised response. Recheck retained facts rather than assuming the correction fixed everything.

**Adapt it:** Use the pattern for an explanation, plan, invitation, or brainstorming session. Choose what a helpful first draft should accomplish and what you will check yourself.

This page's level 1 example is a plain model call: one request and one response, with no search
or tool execution around it. You supply a message and inspect the reply. A chat interface does
not guarantee that architecture: a modern chat product may search, run tools, use memory, or
coordinate several model calls behind one visible reply. The level describes how the task runs,
not the appearance of its message box.

Because there is only one step, the whole outcome depends on what happens inside that one call:
how the model was trained, and what you put in the request. Later levels add retrieval, tools and
loops around the model to make up for what one call alone gets wrong. Before any of that, it
helps to be precise about what "the model" even refers to, because a chat app, the company behind
it, and the model actually answering you are three different things often sharing one name.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 1 · Chat. The same steps are described in the sections below._

## Practical guidance

The name that actually matters is not "Claude" or "ChatGPT," it's the specific model your product
used to answer you. Most chat apps show this in a menu near the message box, a dropdown at the top
of the conversation, or a settings panel; look there before anywhere else. Anthropic's own
documentation names Claude Sonnet 5 as a specific, versioned model, released June 30, 2026 with the
model id `claude-sonnet-5`[1]: that is the kind of name the menu is naming, not the
product's own name.

When a colleague's advice about "Claude" doesn't reproduce, the product name isn't enough
information to debug from. Ask which model their menu showed, not which app they opened: Claude,
Anthropic's own product, is described as "a helpful, intuitive, and powerful collaborator you can
put to work on real tasks"[2], and that description covers more than one model
underneath it. ChatGPT is the same shape: one product name over several possible models. If
neither of you checked the menu, you were never comparing the same thing in the first place.

Two things to check directly in the product itself, not by asking a model about itself from
memory. Type "What model are you, and what is your knowledge cutoff?" and compare the answer
against the menu's own label and the product's own documentation; a model can be wrong about its
own name and date, and the menu is the source that actually decides. And before comparing "Claude"
against another product's newest model by name, check that you are comparing two models rather
than a product against a model, since a product can run more than one and the menu decides which.

The developer and the tool matter mainly when you're building something, not chatting. OpenAI's
own documentation describes its API as what lets a developer "prompt a model and generate text" or
"build agents that use tools and computers"[3]: if a feature was clearly built by
someone else's code calling a model, that's the tool layer, and not something you troubleshoot
from inside a chat window. Ollama states plainly that "Nothing you run locally ever leaves your
machine"[4], which matters only if privacy is the actual question you came with.

None of this matters for a one-off question you can check yourself. It matters when a result
doesn't match what you expected, or what someone else got.

## Implementation details

The Build it example is level 1's entire trick: send the question, change nothing else, and see
what the model does with no documents and no tools available to it. `examples/one_call` asks
about Halvorsen, a fictional appliance maker invented for this site's synthetic documents,
specifically so a model has never legitimately seen its manuals. A correct run at this level
mostly means declining questions it cannot know the answer to, rather than inventing a
plausible-sounding number. That is exactly the failure this level exists to measure, and
exactly what [RAG](/gradient_ascent/techniques/rag/) exists to fix.

The system prompt is the only lever available here, and it says so directly: answer plainly, and
say so plainly when a specific fact is not known instead of guessing at it. There is no chunking,
no search and no schema: one system message, one user message, one call.

`examples/one_call/run.py` (lines 22-39)

```python
def run(question: str, model: Model, embedder: Embedder | None, tracer: Tracer) -> Answer:
    del embedder  # level 1 has no retrieval step
    messages = [
        Message(role="system", content=SYSTEM_PROMPT),
        Message(role="user", content=question),
    ]
    tracer.record(kind="code", decided_by="code", title="Build prompt", detail=question)
    completion = model.complete(messages, max_tokens=400)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask the model",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    return Answer.from_text(completion.text)
```

Two things are worth noticing in the trace. First, calling the model is not itself a model
decision: the code decided to make this call, in this order, before the model said anything: the
`"model"` kind on that step and the `"code"` `decided_by` on the same step answer two different
questions (see `docs/EVALS.md`). Second, there is nothing left to decide once the call returns;
the code does not parse the answer, check it against anything, or call the model again. It hands
back exactly what came back. Run it yourself:

`examples/one_call/README.md` (lines 15-15)

```text
python -m examples.one_call --model stub:scripted
```

A newcomer building their first real thing on top of a model almost always starts here, whether or
not they call it "level 1": one prompt, sent through whichever tool wraps the developer's model,
usually that developer's own API, or a tool like Ollama that can swap which model answers without
changing the calling code.

## When you do not need this

Try [level 0, no model at all](/gradient_ascent/techniques/order-zero/) first if the question has a fixed
vocabulary and repeats often enough that a keyword search or a rule can answer it with no model at
all. Level 0 is fast, free and completely predictable: three things a chat reply cannot
promise.

## Completion in the editor

The editor that finishes your line as you type is this level too, and for people who write code it
is usually the model they touch most hours of the week. GitHub's documentation says "GitHub
Copilot offers coding suggestions as you type"[5]. Cursor describes its own the same
way: "Tab is Cursor's AI-powered autocomplete. It suggests code as you type, based on your recent
edits, surrounding code, and linter errors"[6].

Nobody decides the next step, which is what keeps it on this rung. The editor's code decides when
to ask and what context to send, the model fills in the rest of the line, and you accept it with a
keystroke or keep typing. That is one request and one answer. It is not level 5: a
[coding agent](/gradient_ascent/techniques/coding-agents/) picks each step and decides when it is
finished, and a completion picks nothing.

You do not need it when you already know exactly what the line says, since reading a suggestion
costs more attention than typing eight characters, and at a bench a register write that looks
right is worse than a blank line.

The failure mode follows from that asymmetry: accepting takes one key and checking takes a minute,
so a plausible wrong line lands in the file unread. There is a second thing an unread line can
carry: GitHub's documentation says "GitHub Copilot checks each suggestion for matches with
publicly available code"[7], and that a match is either discarded or offered with a code
reference, depending on a policy setting your account or organization controls[7].

## Failure modes

### Confident answers outside what the model actually knows

- **How to notice it:** The reply is fluent and specific about something the model was never trained on (a fictional product, your own private data, an internal document), instead of saying it does not know.
- **How to test for it:** Ask about something invented for this site's synthetic corpus, like a Halvorsen part number, with no documents attached, and check whether the model declines or guesses.

### No memory beyond what is sent

- **How to notice it:** A follow-up question gets answered as if the earlier part of the conversation never happened, because a single call only sees what is in that one request.
- **How to test for it:** Call the model with only the latest question, no prior turns included, and check whether it can still answer something that depended on earlier context.

### Knowledge cutoff

- **How to notice it:** The model answers confidently about something that changed after its training data ends, using the old fact as if it were current.
- **How to test for it:** Ask about a recent event or a fact you know changed recently, and compare the answer against the model's stated knowledge cutoff.

### No way to check its own answer

- **How to notice it:** Asking "are you sure" is still just another single call; the model may double down or flip its answer with equal confidence either way, since nothing verifies either reply against a source.
- **How to test for it:** Ask the same factual question twice in separate calls, phrased differently, and check whether the two answers actually agree.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one question:** 1
- **Tokens in:** ~40
- **Tokens out:** ~55
- **Wall time:** ~0.6s

**Compared with RAG (level 2).** RAG adds one retrieval step and roughly forty times the input tokens for the same question, in exchange for grounding the answer in real documents instead of whatever the model remembers from training.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

Chat is scored on the same 60-question set as every other level (`docs/EVALS.md`). With no
documents attached, its lookup and numeric scores should sit close to level 0's floor: whatever
it gets right, it gets right from training data alone, which for a fictional appliance maker like
Halvorsen should be close to nothing. The one place a single call can beat a keyword score is
unanswerable questions, if the system prompt's instruction to decline rather than guess actually
holds: it can say "I don't know" in its own words instead of returning an irrelevant passage.

No result file exists yet for any level (see `docs/EVALS.md`). Run
`python scripts/eval_run.py --example one_call --model stub --dry` to project the token cost of a
run before spending anything on a real one.

## Run it

**What to monitor.** The rate of confidently wrong answers on anything outside common knowledge, since a single call has no way to flag its own uncertainty beyond what the prompt asks it to say.

**Cost at volume.** Tokens in track what you send (the question plus any instructions); tokens out track how long the replies run. Both scale linearly with traffic, and there is no retrieval or tool infrastructure running alongside it to add to the bill.

**How it fails in production.** A user asks about something the model was never trained on, or something that changed after its training cutoff, and gets a fluent, wrong answer instead of a refusal. Or an app update quietly drops the instruction to say "I don't know", and nobody notices until a wrong answer causes a real problem.

**What to log.** The full prompt sent (system and user messages), the model id and version, and the raw reply, so a bad answer traces back to what the model was actually asked rather than being guessed at afterward.

## Try it

1. **Use it.** Open a chat app and check which model answered your last message, usually in a menu or settings panel. Ask it its own knowledge cutoff date and compare that against the product's documentation.
2. **Build it.** Run python -m examples.one_call --model stub:scripted from the repo root. With no documents and no tools, the reply declines to give the DR-210's supply voltage and says where the number is: the best answer this level has. Run it again with --model stub: the echo prints the prompt back, the same shape, nothing in it.
3. **Either lane.** Pick a name from the "Out there" list at the bottom of this page: which of the four kinds is it, developer, model, product or tool?
4. **Build it.** At a bench, paste a paragraph from an instrument programming manual into a chat app and ask for a summary: a safe use, since the text is right there to check. Then ask it, with nothing attached, for a specific accuracy figure from memory. A fluent answer is not a reported measurement, and never becomes one.


## Sources

1. [Claude Sonnet 5](https://platform.claude.com/docs/en/models/sonnet-5/overview) — Anthropic, 2026-06-30 (accessed 2026-09-19)
2. [Claude](https://claude.com/product/overview) — Anthropic (accessed 2026-09-19)
3. [OpenAI API Platform Documentation](https://developers.openai.com/api/docs) — OpenAI (accessed 2026-09-19)
4. [Ollama](https://ollama.com/) — Ollama (accessed 2026-09-19)
5. [Getting code suggestions in your IDE with GitHub Copilot](https://docs.github.com/en/copilot/using-github-copilot/getting-code-suggestions-in-your-ide-with-github-copilot) — GitHub (accessed 2026-09-19)
6. [Tab completion](https://cursor.com/docs/tab) — Cursor (accessed 2026-09-19)
7. [Code suggestions](https://docs.github.com/en/copilot/concepts/completions/code-suggestions) — GitHub (accessed 2026-09-19)


Last reviewed 2026-09-19.
