# Reasoning at answer time

_Level 01 · Direct prompting · sourced_

Letting the model think for longer before it answers.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Work through a problem with interacting constraints and inspect the proposed answer against them. The goal is a checkable solution, not a persuasive explanation of how hard the model worked.

**Assumptions:** The constraints must be explicit enough to test. More computation does not establish that the model understood an omitted requirement.

**Design choices:** Use a calculator, search, or solver for parts that have reliable external checks. Extra model effort is useful only if it improves the result enough for the task.

**Request:** Schedule two 45-minute sessions in one room without overlap.

**Starting evidence:** Room opens 10:00. Trainer A leaves 11:00. Trainer B arrives 10:30. Cleanup takes 15 minutes.

**Action and control:** Compare candidate schedules against explicit constraints. These candidates are illustrations, not a model's hidden reasoning.

**Stage records (authored, not executed):**

### Input record

Room opens 10:00. Trainer A leaves 11:00. Trainer B arrives 10:30. Cleanup takes 15 minutes.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use a calculator, search, or solver for parts that have reliable external checks. Extra model effort is useful only if it improves the result enough for the task.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Compare candidate schedules against explicit constraints. These candidates are illustrations, not a model's hidden reasoning.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

A: 10:00–10:45; cleanup to 11:00; B: 11:00–11:45. All supplied constraints satisfied.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

An observable candidate schedule, constraint checker, counterexample, and labeled illustrative cost/quality comparison.

If the result falls short:
When no valid answer is found, distinguish a contradictory specification from a failed attempt. Ask which constraint can change rather than silently relaxing one.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Replace the schedule with a planning or analysis problem. Identify a way to check the answer independently, and compare the extra time against a simpler baseline.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** A: 10:00–10:45; cleanup to 11:00; B: 11:00–11:45. All supplied constraints satisfied.

**Change something — Move Trainer A's arrival to 10:30:** A cannot finish 45 minutes before leaving at 11:00. Report infeasibility rather than invent availability.

**Decision:** Would more computation guarantee a valid schedule now?

**Answer:** No; the constraints may be infeasible.

**Why:** More computation does not ensure correctness; do not present invented hidden reasoning as a real model trace.

**Review criteria:** An observable candidate schedule, constraint checker, counterexample, and labeled illustrative cost/quality comparison.

**Recovery:** When no valid answer is found, distinguish a contradictory specification from a failed attempt. Ask which constraint can change rather than silently relaxing one.

**Adapt it:** Replace the schedule with a planning or analysis problem. Identify a way to check the answer independently, and compare the extra time against a simpler baseline.

Inference-time reasoning is spending more computation after training to get a better answer,
without changing the model itself. It comes in two shapes. The first scales one call: extended
thinking or a reasoning-effort setting lets the model work through a problem before answering, and
Anthropic says that reasoning is billed as output tokens even when the thinking text is not
returned to you[1]. The second scales the number of calls instead: ask the same question
several times and combine the answers, the way self-consistency samples several reasoning paths
and keeps the answer most of them agree on[4].

Anthropic, Google and OpenAI each say close to the same thing about the first kind: more thinking
helps on problems with real intermediate steps and mostly wastes tokens on ones that do not, like
a lookup or a classification[1][2][3]. This page's example uses the
second kind, since it is the one a fixed `StubModel` can demonstrate honestly. Either shape is
still level 1: the model deciding what to think about, or which of several samples to trust, has
not changed who decides what happens next.

This page is sourced, not measured: what thinking longer buys comes from the makers' and
researchers' own papers, and no sampling run here has been scored. It is illustrated.

_The web page for this technique includes an interactive step-through of Level 1 · Self-consistency. The same steps are described in the sections below._

## Practical guidance

Look for a toggle, a slider, or a menu item labeled "thinking," "extended reasoning," or an effort
level from low to high, usually near where you type or in settings. Turn it on, or push it higher,
only for a question with real multiple steps: a word problem with several dependent parts, a plan
that has to account for constraints, a bug you can't spot at a glance. Anthropic's own guidance
describes what that setting buys: a model that visibly works through a problem, restating what's
being asked, trying an approach, checking it, backtracking if it doesn't hold up, before giving a
final answer[1].

For anything else, leave it off or set it low. Ask a plain factual question, such as what year a
law was passed, with the setting on, then again with it off, and time both. If the answer, not
just the wait, comes back identical either way, you've found a question this setting was never
going to help with: the makers' own guidance says to use minimal effort for fact retrieval and
classification, and save the higher settings for coding, math and multi-step
planning[2][3].

Some products never show a toggle at all and instead run several attempts behind the scenes,
showing you only the one they kept. You can't switch that off, but you can still check it: ask the
same real question again in a brand new conversation, worded slightly differently, and see whether
the two answers actually agree. Two confident, different answers to the same question is a sign to
verify the fact independently, not to trust whichever one you saw first.

You'll know the setting earned its cost when turning it on changes the answer on a question you
already know the right answer to, not just when it makes the reply read more thorough. A longer,
more confident-sounding wrong answer is not a win; check it against something you can verify
before trusting the extra length it took to get there.

If an answer is wrong because the model never had a fact it needed, more thinking time will not
fix that, no matter how high the setting goes. That's a missing-information problem, not a
reasoning one, and the fix is giving it the fact directly, not asking it to think harder about the
same gap.

## Implementation details

The example runs self-consistency literally: the same numeric question goes to the model five
times as five independent calls, each reply is asked to end with a line the code can parse
(`Answer: <number>`), and the code returns whichever number the largest share of the five samples
agree on.

`examples/inference_time_reasoning/run.py` (lines 34-59)

```python
def run(question: str, model: Model, tracer: Tracer, *, n: int = N_SAMPLES) -> Answer:
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=question)]
    tracer.record(kind="code", decided_by="code", title="Build one fixed prompt", detail=question)
    votes: Counter[str] = Counter()
    for i in range(n):
        completion = model.complete(messages, max_tokens=200)
        answer = _extract(completion.text) or "no answer"
        votes[answer] += 1
        tracer.record(
            kind="model",
            decided_by="code",
            title=f"Sample {i + 1} of {n}",
            detail=completion.text[:200],
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
    winner, count = votes.most_common(1)[0]
    tracer.record(kind="code", decided_by="code", title="Tally the votes", detail=f"{dict(votes)}")
    tracer.record(
        kind="code",
        decided_by="code",
        title="Return the majority answer",
        detail=f"{winner} ({count}/{n} samples agreed)",
    )
    return Answer(text=f"{winner} ({count}/{n} samples agreed)", citations=[])
```

Extracting the final number from free-form reasoning text is its own small, fixed piece of code
(`_extract`): read from the bottom for a line starting `Answer:` and pull the number out of it, so
a sample that reasons at length still ends in something machine-checkable. Five samples from a
scripted stub exercise the vote itself, not a claim about how often real models agree with
themselves on a hard question: that claim (whether five samples on a real model land on the
right number more often than one sample does) is exactly what this site's eval set is built to
measure, once a real run exists (see `docs/EVALS.md`).

Run it yourself:

`examples/inference_time_reasoning/README.md` (lines 14-14)

```text
python -m examples.inference_time_reasoning --model stub:scripted
```

Every step is `decided_by: "code"`: the sample count is fixed, and the code always takes the
majority regardless of what any individual sample said. A model choosing to think longer inside
one call (the other shape of this technique) would still be `decided_by: "code"` by this site's
definition too: the code decided to turn thinking on or set an effort level, and the model
deciding what to think about is not the same as the model deciding what the control flow does
next (see `docs/EVALS.md`).

## When you do not need this

Skip the extra tokens if a single plain call already gets the answer right on repeat tries: test
that before assuming more thinking or more samples will help. And if the model is wrong because it
was never given a fact it needed, not because it reasoned badly, more reasoning effort does not
fix that; [RAG](/gradient_ascent/techniques/rag/) or a better prompt fixes a missing-information
problem, not a reasoning one.

## Failure modes

### A systematic error looks unanimous

- **How to notice it:** All samples make the same mistake (a shared misreading of the question, an arithmetic slip everyone reproduces), so the majority vote reports high confidence in a wrong answer.
- **How to test for it:** Check a case where the correct answer is already known, and verify the votes are not unanimous for a wrong one.

### No answer to extract

- **How to notice it:** A sample reasons at length but never states its answer in the expected format, so it silently falls into "no answer" instead of being flagged as a parsing failure.
- **How to test for it:** Check the "no answer" bucket's share of votes across a batch of runs, not just which answer won.

### Reasoning tokens with nothing to reason about

- **How to notice it:** Turning on extended thinking or a high effort level for a simple lookup or classification burns tokens and adds latency with no change in the answer.
- **How to test for it:** Compare token count and wall time with thinking on versus off on the same simple question, holding the question fixed.

### A near-tie decided arbitrarily

- **How to notice it:** The votes split close to evenly and the code picks whichever answer happened to be tallied first, presenting it with the same confidence as a clear majority.
- **How to test for it:** Log the full vote distribution, not just the winner, and treat a close vote differently from a landslide.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one question:** 5
- **Tokens in (total):** ~260
- **Tokens out (total):** ~450
- **Wall time:** ~2s

**Compared with chat (level 1).** Five independent samples cost roughly five times a single call in tokens, and roughly five times the wall time run one after another (or close to one call's wall time if run in parallel, at the same total token cost), for a better chance at a correct answer on questions with more than one path to it.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

Self-consistency and extended thinking are graded like every technique here: scored against the
same 60-question set (`docs/EVALS.md`), with the sample count or the effort level recorded as a
setting on the run rather than a fixed part of the technique. The comparison that actually matters
is one sample against five, or low effort against high, on the exact same questions, since the
whole claim is "more inference-time computation raises the score for some class of model, on some
kinds of question", and the site's claim rule requires naming which model class and which kinds
that holds for once a result file exists.

This example samples one arithmetic question five times and returns the majority answer. It
reads no documents and cites nothing, so `scripts/eval_run.py` will not score it against the
60-question set and says so instead of returning a number measured on the wrong task
(`docs/EVALS.md`). Measuring it needs questions with a checkable numeric answer, run at one
sample, three and five: the score at each setting, the share of runs where the samples agreed,
and the token cost of each, since five samples cost about five times one.

## Run it

**What to monitor.** Agreement rate across samples (unanimous versus split), tracked separately from raw accuracy. A model that agrees with itself confidently and is wrong needs a different fix than one that disagrees with itself but is right on the majority side.

**Cost at volume.** Cost multiplies by the sample count, or by the extra reasoning tokens for a single deeper call, on every question whether or not that question needed it. This is the one technique on the site where the multiplier is a number you choose directly.

**How it fails in production.** A question type that used to have one dominant right answer starts splitting votes evenly after a data or prompt change, and the majority pick becomes close to a coin flip without the interface showing any less confidence than before.

**What to log.** Every sample's raw text and extracted answer, the full vote tally, and which one won, so a bad final answer can be told apart from a bad extraction of an otherwise fine sample.

## Try it

1. **Use it.** Find a "thinking" or "reasoning effort" toggle in a chat app. Ask a simple factual question with it on and off, compare the wait and the answer, then a multi-step problem.
2. **Build it.** Run python -m examples.inference_time_reasoning --model stub:scripted from the repo root: five samples, three agreeing on 79.50, two wrong, and the vote picking the right one. Now edit SCRIPTED (examples/inference_time_reasoning/__main__.py) so all five answers differ: the vote still returns one, reported as 1/5 agreed. The tally, not the answer, says how far to trust it.
3. **Either lane.** Pick a question a model got wrong. Would thinking longer have fixed it, or did it lack the information?


## Sources

1. [Thinking](https://platform.claude.com/docs/en/build-with-claude/thinking) — Anthropic (accessed 2026-09-19)
2. [Thinking](https://ai.google.dev/gemini-api/docs/thinking) — Google (accessed 2026-09-19)
3. [Reasoning models](https://developers.openai.com/api/docs/guides/reasoning) — OpenAI (accessed 2026-09-19)
4. [Self-Consistency Improves Chain of Thought Reasoning in Language Models](https://arxiv.org/abs/2203.11171) — arXiv (Google Research, UC Santa Barbara), 2022-03-21 (accessed 2026-09-19)


Last reviewed 2026-09-19.
