Level 01 · Direct prompting

Reasoning at answer time

Letting the model think for longer before it answers.

Sourced

Concept at a glance

Spend more work on the answer before returning it.

SequenceConceptual illustration
Spend more work on the answer before returning it.Hard question leads to Reason or compare. Reason or compare leads to Final answer. Extra reasoning happens while answering; it does not train a new model.Hard questionA task worth extra effortReason or compareMore work at answer timeFinal answerStill needs checkingSpend more work on the answer before returning it.Hard question leads to Reason or compare. Reason or compare leads to Final answer. Extra reasoning happens while answering; it does not train a new model.Hard questionA task worth extra effortReason or compareMore work at answer timeFinal answerStill needs checking
Read the connections in words
  • Hard question → Reason or compare: More work at answer time.
  • Reason or compare → Final answer: Still needs checking.
Key idea

Extra reasoning happens while answering; it does not train a new model.

A focused everyday life example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Reasoning at answer time: see it in practice.

Allocating additional inference computation to work through a problem before returning an answer.

What you’ll walk through

Work through a problem with interacting constraints and inspect the proposed answer against them. The goal is a checkable solution, not a persuasive explanation of how hard the model worked.

The task in this version

Schedule two 45-minute sessions in one room without overlap.

What you’ll learn to check

An observable candidate schedule, constraint checker, counterexample, and labeled illustrative cost/quality comparison.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Everyday lifeAn authored case with its own evidence, changed condition, and decision.
The task in this example

Schedule two 45-minute sessions in one room without overlap.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Room opens 10:00. Trainer A leaves 11:00. Trainer B arrives 10:30. Cleanup takes 15 minutes.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

The constraints must be explicit enough to test. More computation does not establish that the model understood an omitted requirement.

1 / 6

Apply this to your project

Describe your task to your own model and use Reasoning at answer time as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Inference-time reasoning is spending more computation after training to get a better answer, without changing the model itself. It comes in two shapes. The first scales one call: extended thinking or a reasoning-effort setting lets the model work through a problem before answering, and Anthropic says that reasoning is billed as output tokens even when the thinking text is not returned to you[1]. The second scales the number of calls instead: ask the same question several times and combine the answers, the way self-consistency samples several reasoning paths and keeps the answer most of them agree on[4].

Anthropic, Google and OpenAI each say close to the same thing about the first kind: more thinking helps on problems with real intermediate steps and mostly wastes tokens on ones that do not, like a lookup or a classification[1][2][3]. This page’s example uses the second kind, since it is the one a fixed StubModel can demonstrate honestly. Either shape is still level 1: the model deciding what to think about, or which of several samples to trust, has not changed who decides what happens next.

This page is sourced, not measured: what thinking longer buys comes from the makers’ and researchers’ own papers, and no sampling run here has been scored. It is illustrated.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Self-consistency: sample five times, vote

Ask the same numeric question five independent times and keep the answer most samples agree on.

Level 1 · Direct prompting, sampled 5×
QuestionQuestionBuild one fixed promptBuild one fixed promptMODEL5 independent samples5 independent samplesTally votes, take the majorityTally votes,take the majorityAnswerAnswerQuestionQuestionBuild one fixed promptBuild one fixed promptMODEL5 independent samples5 independent samplesTally votes, take the majorityTally votes,take the majorityAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 08Your code chose

The question arrives

"What is the total price to replace the heating elements
on both a DW-300 and a DW-480?"
0 tokens · 0 ms

Practical guidance

Look for a toggle, a slider, or a menu item labeled “thinking,” “extended reasoning,” or an effort level from low to high, usually near where you type or in settings. Turn it on, or push it higher, only for a question with real multiple steps: a word problem with several dependent parts, a plan that has to account for constraints, a bug you can’t spot at a glance. Anthropic’s own guidance describes what that setting buys: a model that visibly works through a problem, restating what’s being asked, trying an approach, checking it, backtracking if it doesn’t hold up, before giving a final answer[1].

For anything else, leave it off or set it low. Ask a plain factual question, such as what year a law was passed, with the setting on, then again with it off, and time both. If the answer, not just the wait, comes back identical either way, you’ve found a question this setting was never going to help with: the makers’ own guidance says to use minimal effort for fact retrieval and classification, and save the higher settings for coding, math and multi-step planning[2][3].

Some products never show a toggle at all and instead run several attempts behind the scenes, showing you only the one they kept. You can’t switch that off, but you can still check it: ask the same real question again in a brand new conversation, worded slightly differently, and see whether the two answers actually agree. Two confident, different answers to the same question is a sign to verify the fact independently, not to trust whichever one you saw first.

You’ll know the setting earned its cost when turning it on changes the answer on a question you already know the right answer to, not just when it makes the reply read more thorough. A longer, more confident-sounding wrong answer is not a win; check it against something you can verify before trusting the extra length it took to get there.

If an answer is wrong because the model never had a fact it needed, more thinking time will not fix that, no matter how high the setting goes. That’s a missing-information problem, not a reasoning one, and the fix is giving it the fact directly, not asking it to think harder about the same gap.

Implementation details

The example runs self-consistency literally: the same numeric question goes to the model five times as five independent calls, each reply is asked to end with a line the code can parse (Answer: <number>), and the code returns whichever number the largest share of the five samples agree on.

examples/inference_time_reasoning/run.py · lines 34–59
def run(question: str, model: Model, tracer: Tracer, *, n: int = N_SAMPLES) -> Answer:
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=question)]
    tracer.record(kind="code", decided_by="code", title="Build one fixed prompt", detail=question)
    votes: Counter[str] = Counter()
    for i in range(n):
        completion = model.complete(messages, max_tokens=200)
        answer = _extract(completion.text) or "no answer"
        votes[answer] += 1
        tracer.record(
            kind="model",
            decided_by="code",
            title=f"Sample {i + 1} of {n}",
            detail=completion.text[:200],
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
    winner, count = votes.most_common(1)[0]
    tracer.record(kind="code", decided_by="code", title="Tally the votes", detail=f"{dict(votes)}")
    tracer.record(
        kind="code",
        decided_by="code",
        title="Return the majority answer",
        detail=f"{winner} ({count}/{n} samples agreed)",
    )
    return Answer(text=f"{winner} ({count}/{n} samples agreed)", citations=[])

Extracting the final number from free-form reasoning text is its own small, fixed piece of code (_extract): read from the bottom for a line starting Answer: and pull the number out of it, so a sample that reasons at length still ends in something machine-checkable. Five samples from a scripted stub exercise the vote itself, not a claim about how often real models agree with themselves on a hard question: that claim (whether five samples on a real model land on the right number more often than one sample does) is exactly what this site’s eval set is built to measure, once a real run exists (see docs/EVALS.md).

Run it yourself:

examples/inference_time_reasoning/README.md · lines 14–14
python -m examples.inference_time_reasoning --model stub:scripted

Every step is decided_by: "code": the sample count is fixed, and the code always takes the majority regardless of what any individual sample said. A model choosing to think longer inside one call (the other shape of this technique) would still be decided_by: "code" by this site’s definition too: the code decided to turn thinking on or set an effort level, and the model deciding what to think about is not the same as the model deciding what the control flow does next (see docs/EVALS.md).

When you do not need this

Skip the extra tokens if a single plain call already gets the answer right on repeat tries: test that before assuming more thinking or more samples will help. And if the model is wrong because it was never given a fact it needed, not because it reasoned badly, more reasoning effort does not fix that; RAG or a better prompt fixes a missing-information problem, not a reasoning one.

Failure modes

A systematic error looks unanimous

How to notice it
All samples make the same mistake (a shared misreading of the question, an arithmetic slip everyone reproduces), so the majority vote reports high confidence in a wrong answer.
How to test for it
Check a case where the correct answer is already known, and verify the votes are not unanimous for a wrong one.

No answer to extract

How to notice it
A sample reasons at length but never states its answer in the expected format, so it silently falls into "no answer" instead of being flagged as a parsing failure.
How to test for it
Check the "no answer" bucket's share of votes across a batch of runs, not just which answer won.

Reasoning tokens with nothing to reason about

How to notice it
Turning on extended thinking or a high effort level for a simple lookup or classification burns tokens and adds latency with no change in the answer.
How to test for it
Compare token count and wall time with thinking on versus off on the same simple question, holding the question fixed.

A near-tie decided arbitrarily

How to notice it
The votes split close to evenly and the code picks whichever answer happened to be tallied first, presenting it with the same confidence as a clear majority.
How to test for it
Log the full vote distribution, not just the winner, and treat a close vote differently from a landslide.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

5Model calls, one question
~260Tokens in (total)
~450Tokens out (total)
~2sWall time
Compared with chat (level 1)Five independent samples cost roughly five times a single call in tokens, and roughly five times the wall time run one after another (or close to one call's wall time if run in parallel, at the same total token cost), for a better chance at a correct answer on questions with more than one path to it.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

Self-consistency and extended thinking are graded like every technique here: scored against the same 60-question set (docs/EVALS.md), with the sample count or the effort level recorded as a setting on the run rather than a fixed part of the technique. The comparison that actually matters is one sample against five, or low effort against high, on the exact same questions, since the whole claim is “more inference-time computation raises the score for some class of model, on some kinds of question”, and the site’s claim rule requires naming which model class and which kinds that holds for once a result file exists.

This example samples one arithmetic question five times and returns the majority answer. It reads no documents and cites nothing, so scripts/eval_run.py will not score it against the 60-question set and says so instead of returning a number measured on the wrong task (docs/EVALS.md). Measuring it needs questions with a checkable numeric answer, run at one sample, three and five: the score at each setting, the share of runs where the samples agreed, and the token cost of each, since five samples cost about five times one.

Run it

What to monitor

Agreement rate across samples (unanimous versus split), tracked separately from raw accuracy. A model that agrees with itself confidently and is wrong needs a different fix than one that disagrees with itself but is right on the majority side.

Cost at volume

Cost multiplies by the sample count, or by the extra reasoning tokens for a single deeper call, on every question whether or not that question needed it. This is the one technique on the site where the multiplier is a number you choose directly.

How it fails in production

A question type that used to have one dominant right answer starts splitting votes evenly after a data or prompt change, and the majority pick becomes close to a coin flip without the interface showing any less confidence than before.

What to log

Every sample's raw text and extracted answer, the full vote tally, and which one won, so a bad final answer can be told apart from a bad extraction of an otherwise fine sample.

Try it

  1. Use it

    Find a "thinking" or "reasoning effort" toggle in a chat app. Ask a simple factual question with it on and off, compare the wait and the answer, then a multi-step problem.

  2. Build it

    Run python -m examples.inference_time_reasoning --model stub:scripted from the repo root: five samples, three agreeing on 79.50, two wrong, and the vote picking the right one. Now edit SCRIPTED (examples/inference_time_reasoning/__main__.py) so all five answers differ: the vote still returns one, reported as 1/5 agreed. The tally, not the answer, says how far to trust it.

  3. Either lane

    Pick a question a model got wrong. Would thinking longer have fixed it, or did it lack the information?

How it connects

Before, after and instead of this

Move up when

  • Write and checkThe same mistake shows up across every sample or every extra round of thinking, so more computation on the same approach stops helping and the draft needs checking against a stated criterion instead.
Optional: products, tools, and models

5 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Work through a difficult calculation

Give the model more reasoning effort or compare several candidate solutions before accepting an answer.

Out there

Named products, tools and models

Models6
  • Claude Opus 5Anthropic · frontier model
  • DeepSeek V4DeepSeek · open-weight modelSuperseded by DeepSeek-V4.1-Flash
  • DeepSeek-V4.1-FlashDeepSeek · open-weight model
  • Gemini 3.1 ProGoogle · frontier model
  • GPT-6 AstraOpenAI · frontier model
  • Grok 4.6SpaceXAI · frontier model

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Thinking · Anthropic (accessed 09/19/2026)
  2. Thinking · Google (accessed 09/19/2026)
  3. Reasoning models · OpenAI (accessed 09/19/2026)
  4. Self-Consistency Improves Chain of Thought Reasoning in Language Models · arXiv (Google Research, UC Santa Barbara), 03/21/2022 (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page