Topics at every level

Calibrating trust

Learning, from results over time, how much to rely on a model without checking.

Sourced

Concept at a glance

Calibrate reliance from observed results.

Feedback loopConceptual illustration
Calibrate reliance from observed results.A bounded task leads to Model does work. Model does work leads to Verify results. Verify results leads to Adjust reliance. Adjust reliance leads to Model does work as feedback. Trust should be specific to a task and revised when evidence changes.A bounded taskSet the stakes and checksModel does workWithin that scopeVerify resultsRecord successes and missesAdjust relianceChange scope or review depthCalibrate reliance from observed results.A bounded task leads to Model does work. Model does work leads to Verify results. Verify results leads to Adjust reliance. Adjust reliance leads to Model does work as feedback. Trust should be specific to a task and revised when evidence changes.A bounded taskSet the stakes and checksModel does workWithin that scopeVerify resultsRecord successes and missesAdjust relianceChange scope or review depth

Ending or continuingRevisit reliance when the task, model, or evidence changes.

Read the connections in words
  • A bounded task → Model does work: Within that scope.
  • Model does work → Verify results: Record successes and misses.
  • Verify results → Adjust reliance: Change scope or review depth.
  • Adjust reliance → Model does work: feedback informs another turn.
Key idea

Trust should be specific to a task and revised when evidence changes.

CHOOSE YOUR PERSPECTIVE

Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Calibrating trust: see it in practice.

Calibrating reliance on a model from observed performance in a particular task and context.

What you’ll walk through

Follow a sequence of outcomes into a decision about how much oversight to use next. Inspect whether evidence of reliability transfers to the task now being attempted.

The task in this version

Decide how closely to review an assistant's event plans.

What you’ll learn to check

A small labeled performance history, per-task review policy, a novel-case failure, and a justified change in oversight.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Everyday lifeAn authored case with its own evidence, changed condition, and decision.
The task in this example

Decide how closely to review an assistant's event plans.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Routine drafts worked in a small reviewed sample; no history with accessibility or contracts.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Success on familiar easy cases may not transfer to new contexts. Confidence should concern a specific capability under specific conditions.

1 / 6

Apply this to your project

Describe your task to your own model and use Calibrating trust as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Calibrating trust means keeping how much you rely on a model without checking in line with how often it has actually been right on tasks like the one in front of you. It is a record, not an impression, and it is built one checked result at a time. That is why reviewing comes first: the reviews are the entries. It is the last of the four skills on the operator craft topic, and what the record is for is the next round of delegating.

Both directions cost something. Over-trust looks exactly like things going well, right up until the wrong answer that mattered goes through unread. Under-trust looks like diligence: re-doing work a model has done reliably a hundred times, or avoiding a task it would genuinely help with because a different task went badly once. Neither shows up as an error anywhere.

The unit of trust is a task type, not a model. Reliably right at summarizing a document you can check says nothing about arithmetic, or about a question outside anything it was trained on. A single global “I trust this one” hides which specific things it has earned.

This page is sourced, not measured: what the makers claim below is quoted from their own pages, and no claim here has been checked against a run of this site’s own.

Practical guidance

A record for one person is five columns, takes about twenty seconds a row, and lives wherever you already keep notes:

Date Task type What you checked Right? What you changed
3/12/2027 renewal dates from a contract all 4 dates against the clauses yes nothing
3/12/2027 plain-language summary of a policy the 2 exclusions it listed no added the third exclusion

Task type keeps the record usable later: separate lines even when the same product did both. What you changed is the check itself: nothing, a word, or the whole thing. After a month, read down that column for one task type. Mostly “nothing” means you can safely sample instead of reading every one; anything else means you aren’t ready to stop checking, whatever the product’s reputation is. Write the number down; it’s next month’s sampling rate, not a feeling you’ll remember correctly later.

Two rules keep it honest. Log the checks that came back fine, not only the corrections, or the record reads like a catalog of disasters. And start a fresh page whenever what you’re trusting changes: a new model version, an edited prompt, a different tool. Anthropic says exactly this about its own published techniques: where one names a specific model, “treat it as measured on that model and re-check it against your own evals before applying it to another.”[1] A record built on last quarter’s version doesn’t transfer just because the product name did.

Skip it for a one-off task you’ll never ask again, or something so low-stakes a wrong answer costs nothing to fix. Keep it for whatever you catch yourself about to trust from memory instead of a count.

Be wary of research that sounds like it settles this. A 2019 complacency scale was built on Mechanical Turk respondents whose “experience with automation was predominantly with relatively low-stakes and common forms of automation, such as in-car navigation systems,” which its own authors name as a limitation[2]; those participants weren’t supervising a model at work, so treat the finding as a reason to keep your own count, not a number about your job. A 2025 survey defines over-reliance as “relying on LLMs beyond their capabilities”[3] and argues for measurement over impression, which is exactly what the table above is.

Implementation details

A team cannot keep one person’s notebook. What replaces it is a fixed set of checkable questions, scored the same way every time and kept as a file: the evals topic pointed at a running system rather than at a prompt being drafted. “We have been using it a while and it seems fine” is the thing the file exists to replace.

Three properties make a team record worth keeping:

  • Broken out by task kind, the way this site’s own eval set splits its questions, so a good score on easy lookups cannot stand in for the record on the cases nobody re-checked.
  • Versioned by what produced it: model id, prompt version, tool set. A rise or a fall is only informative if you can name what changed; without that, the record is a mood.
  • Fed automatically where it can be. A result that carries a citation, the way RAG’s does, can have the mechanical part checked by machine: the reviewing page’s own checker reports whether a figure the answer states appears in the section it cites. That is one row of evidence per answer without a person reading the whole thing, and it is a presence check, so it is a floor under the record, not the record.

Sample by hand on top of that, at a rate you write down. The machine check and the human sample answer different questions, and an automated number rising while nobody has read an output in six weeks is exactly the state that looks safest and is not.

One design decision belongs here rather than in the record: what the system does when its own confidence is low. A system that can say “I did not find this” gives a reviewer a signal worth logging; one that always produces an answer makes every result look identical from outside, and a record over identical-looking results is much more expensive to keep.

At each level

  • Conventional software: nothing to calibrate: a rule passes its tests or it does not, and it behaves the same way tomorrow.
  • Direct prompting: every reply is independent, so the record is simply a tally per task type with nothing else to attribute a change to.
  • Added context: the record has to separate “found the right source and read it wrong” from “never found it”, because those two have different fixes.
  • Workflows: the steps are fixed, so trust can be tracked per step, and one unreliable stage stops dragging down the ones around it.
  • Tool use: the record now needs to cover the action taken, not only the sentence produced: a wrong call can be reported in perfectly correct prose.
  • Agent loops: the unit becomes a whole run of variable length, so the record tracks outcomes and cost per run rather than accuracy per answer.
  • Teams of Agents: agreement between agents is not evidence. The record has to cover the division of labour, since a team can be confidently wrong together.
  • Always-on agents: the record is the only thing standing in for someone watching, so it has to be written by the system itself and read by a person on a schedule.

Practices

  • Keep the record per task type. “It has been good lately” is not a record.
  • Log the checks that came back fine as well as the ones that did not; otherwise the record only contains disasters and reads like one.
  • Set the sampling rate from the last month’s edit rate, and write the number down where someone else can see it.
  • Re-check after any change to the model, the prompt or the tools, and mark the record with what changed.
  • Name the tasks you deliberately do not trust, and what evidence would change that. An untested assumption in the cautious direction is still an untested assumption.

Run it

What to monitor

Per task type, the edit rate (how often a result is accepted unchanged) next to the sampling rate actually being used. Those two numbers moving apart is the whole subject of this page, in either direction.

Cost at volume

Keeping the record costs roughly the same per checked result at any volume, so the sampling rate, not the traffic, sets the bill. What volume changes is the cost of being miscalibrated: the same error rate is a nuisance at ten results a day and a recall at ten thousand.

How it fails in production

The record stops being written before it stops being cited. Months later a decision is justified with 'it has been reliable', and the last entry anyone made was before two model upgrades and a prompt rewrite.

What to log

Every checked result against what was actually correct, tagged with the task type and with the model and prompt version that produced it. Untagged accuracy cannot be compared across a change, which is the only comparison that matters.

Try it

  1. Use it

    Pick a task you now let a model do without checking. Write down when you last verified one of its answers and what you found. If you cannot remember, that gap is the distance between your trust and your record, and the next five results are the cheapest rows you will ever add.

  2. Build it

    Take one output type from a system you use or built. Sketch the automatic check for it (what would a machine compare against what) and say what that check would NOT catch. The second half is what the human sample is for.

  3. Either lane

    Name one task you trust a model on and one you deliberately do not. For each, write the evidence the position rests on. If either answer is a feeling rather than a count, that is the one to start a record for.

How it connects

Before, after and instead of this

Often used with

Optional: products, tools, and models

Concrete examples

In practice

Adjust review depth from evidence

Track how a model performs on a specific recurring task before deciding which parts can receive lighter review.

An illustrative task example. No verified product or tool is currently listed for this concept.

Out there

Named products, tools and models

No product, tool or model is registered against this page yet. The names index lists every one the site does name, and which technique each belongs to.

Open the names index →

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Prompting best practices · Anthropic (Claude Platform Docs) (accessed 09/19/2026)
  2. Automation-Induced Complacency Potential: Development and Validation of a New Scale · Frontiers in Psychology (Merritt et al.), 02/19/2019 (accessed 09/19/2026)
  3. Measuring and mitigating overreliance to build human-compatible AI · arXiv (Ibrahim et al.), 09/08/2025 (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page