# Calibrating trust

_Topics at every level · sourced_

Learning, from results over time, how much to rely on a model without checking.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a sequence of outcomes into a decision about how much oversight to use next. Inspect whether evidence of reliability transfers to the task now being attempted.

**Assumptions:** Success on familiar easy cases may not transfer to new contexts. Confidence should concern a specific capability under specific conditions.

**Design choices:** Increase autonomy gradually where observed performance and recoverability support it. Keep direct checks on consequential or unfamiliar outputs.

**Request:** Decide how closely to review an assistant's event plans.

**Starting evidence:** Routine drafts worked in a small reviewed sample; no history with accessibility or contracts.

**Action and control:** Base oversight on task-specific evidence and consequences.

**Stage records (authored, not executed):**

### Input record

Routine drafts worked in a small reviewed sample; no history with accessibility or contracts.

What changed: Establish the facts supplied for this version of the task.

### Design note

Increase autonomy gradually where observed performance and recoverability support it. Keep direct checks on consequential or unfamiliar outputs.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Base oversight on task-specific evidence and consequences.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Light checks for routine wording; close source review for novel accessibility and contract claims.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

A small labeled performance history, per-task review policy, a novel-case failure, and a justified change in oversight.

If the result falls short:
When a new failure appears, narrow reliance and investigate its conditions. Neither one success nor one mistake establishes universal trustworthiness.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for your own assistant or a team service. Track the tasks it handles reliably and the situations that still need closer review.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Light checks for routine wording; close source review for novel accessibility and contract claims.

**Change something — Assistant expresses high confidence:** Confidence supplies no relevant performance evidence. Keep task-appropriate review.

**Decision:** Should confident language reduce oversight on unfamiliar work?

**Answer:** No; use evidence and task risk.

**Why:** Success on routine cases does not establish reliability on unusual ones; confidence and polished language are weak evidence.

**Review criteria:** A small labeled performance history, per-task review policy, a novel-case failure, and a justified change in oversight.

**Recovery:** When a new failure appears, narrow reliance and investigate its conditions. Neither one success nor one mistake establishes universal trustworthiness.

**Adapt it:** Use this for your own assistant or a team service. Track the tasks it handles reliably and the situations that still need closer review.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a sequence of outcomes into a decision about how much oversight to use next. Inspect whether evidence of reliability transfers to the task now being attempted.

**Assumptions:** Success on familiar easy cases may not transfer to new contexts. Confidence should concern a specific capability under specific conditions.

**Design choices:** Increase autonomy gradually where observed performance and recoverability support it. Keep direct checks on consequential or unfamiliar outputs.

**Request:** Choose how closely to review an assistant's new test-project code.

**Starting evidence:** Prior work followed file conventions reliably. No validated history with a new instrument or timing-sensitive measurement.

**Action and control:** Calibrate reliance by task and evidence instead of transferring trust from formatting to measurement correctness.

**Stage records (authored, not executed):**

### Input record

Prior work followed file conventions reliably. No validated history with a new instrument or timing-sensitive measurement.

What changed: Establish the facts supplied for this version of the task.

### Design note

Increase autonomy gradually where observed performance and recoverability support it. Keep direct checks on consequential or unfamiliar outputs.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Calibrate reliance by task and evidence instead of transferring trust from formatting to measurement correctness.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Reuse scaffolding with review; scrutinize unfamiliar driver calls and validate timing with the approved human-led process.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Separate conventional file checks from measurement validation and review each appropriately.

If the result falls short:
When a new failure appears, narrow reliance and investigate its conditions. Neither one success nor one mistake establishes universal trustworthiness.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for your own assistant or a team service. Track the tasks it handles reliably and the situations that still need closer review.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Reuse scaffolding with review; scrutinize unfamiliar driver calls and validate timing with the approved human-led process.

**Change something — Agent says it is certain the new timing is correct:** Confidence supplies no timing evidence. Retain verification appropriate to the new behavior.

**Decision:** Does reliability on scaffolding establish reliability on hardware timing?

**Answer:** No; these require different evidence.

**Why:** Trust should follow demonstrated capabilities and consequences of error.

**Review criteria:** Separate conventional file checks from measurement validation and review each appropriately.

**Recovery:** When a new failure appears, narrow reliance and investigate its conditions. Neither one success nor one mistake establishes universal trustworthiness.

**Adapt it:** Use this for your own assistant or a team service. Track the tasks it handles reliably and the situations that still need closer review.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a sequence of outcomes into a decision about how much oversight to use next. Inspect whether evidence of reliability transfers to the task now being attempted.

**Assumptions:** Success on familiar easy cases may not transfer to new contexts. Confidence should concern a specific capability under specific conditions.

**Design choices:** Increase autonomy gradually where observed performance and recoverability support it. Keep direct checks on consequential or unfamiliar outputs.

**Request:** Decide whether to reduce review on recurring weekly reports.

**Starting evidence:** Prior ten reports were correct for two stable projects. Three new projects use different source systems.

**Action and control:** Treat the new sources and project definitions as a changed operating context.

**Stage records (authored, not executed):**

### Input record

Prior ten reports were correct for two stable projects. Three new projects use different source systems.

What changed: Establish the facts supplied for this version of the task.

### Design note

Increase autonomy gradually where observed performance and recoverability support it. Keep direct checks on consequential or unfamiliar outputs.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Treat the new sources and project definitions as a changed operating context.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Keep review on new project sections until evidence supports their reliability; routine history remains relevant only within its scope.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Track errors by project/source and document why review effort changes.

If the result falls short:
When a new failure appears, narrow reliance and investigate its conditions. Neither one success nor one mistake establishes universal trustworthiness.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for your own assistant or a team service. Track the tasks it handles reliably and the situations that still need closer review.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Keep review on new project sections until evidence supports their reliability; routine history remains relevant only within its scope.

**Change something — Assistant produces the new sections with polished certainty:** Style does not demonstrate correct source interpretation. Inspect fresh evidence and unresolved conflicts.

**Decision:** Should routine success remove review for unfamiliar sources?

**Answer:** No; verify the new context first.

**Why:** Reliance needs task- and source-specific evidence rather than global confidence.

**Review criteria:** Track errors by project/source and document why review effort changes.

**Recovery:** When a new failure appears, narrow reliance and investigate its conditions. Neither one success nor one mistake establishes universal trustworthiness.

**Adapt it:** Use this for your own assistant or a team service. Track the tasks it handles reliably and the situations that still need closer review.

Calibrating trust means keeping how much you rely on a model without checking in line with how
often it has actually been right on tasks like the one in front of you. It is a record, not an
impression, and it is built one checked result at a time. That is why
[reviewing](/gradient_ascent/techniques/reviewing/) comes first: the reviews are the entries.
It is the last of the four skills on the [operator
craft](/gradient_ascent/techniques/operator-craft/) topic, and what the record is *for* is the next round of
[delegating](/gradient_ascent/techniques/delegating/).

Both directions cost something. Over-trust looks exactly like things going well, right up until
the wrong answer that mattered goes through unread. Under-trust looks like diligence: re-doing
work a model has done reliably a hundred times, or avoiding a task it would genuinely help with
because a different task went badly once. Neither shows up as an error anywhere.

The unit of trust is a task type, not a model. Reliably right at summarizing a document you can
check says nothing about arithmetic, or about a question outside anything it was trained on. A
single global "I trust this one" hides which specific things it has earned.

This page is sourced, not measured: what the makers claim below is quoted from their own pages,
and no claim here has been checked against a run of this site's own.

## Practical guidance

A record for one person is five columns, takes about twenty seconds a row, and lives wherever
you already keep notes:

| Date | Task type | What you checked | Right? | What you changed |
|---|---|---|---|---|
| 3/12/2027 | renewal dates from a contract | all 4 dates against the clauses | yes | nothing |
| 3/12/2027 | plain-language summary of a policy | the 2 exclusions it listed | no | added the third exclusion |

**Task type** keeps the record usable later: separate lines even when the same product did both.
**What you changed** is the check itself: nothing, a word, or the whole thing. After a month,
read down that column for one task type. Mostly "nothing" means you can safely sample instead of
reading every one; anything else means you aren't ready to stop checking, whatever the product's
reputation is. Write the number down; it's next month's sampling rate, not a feeling you'll
remember correctly later.

Two rules keep it honest. Log the checks that came back fine, not only the corrections, or the
record reads like a catalog of disasters. And start a fresh page whenever what you're trusting
changes: a new model version, an edited prompt, a different tool. Anthropic says exactly this
about its own published techniques: where one names a specific model, "treat it as measured on
that model and re-check it against your own evals before applying it to another."[1] A
record built on last quarter's version doesn't transfer just because the product name did.

Skip it for a one-off task you'll never ask again, or something so low-stakes a wrong answer
costs nothing to fix. Keep it for whatever you catch yourself about to trust from memory instead
of a count.

Be wary of research that sounds like it settles this. A 2019 complacency scale was built on
Mechanical Turk respondents whose "experience with automation was predominantly with relatively
low-stakes and common forms of automation, such as in-car navigation systems," which its own
authors name as a limitation[2]; those participants weren't supervising a model at
work, so treat the finding as a reason to keep your own count, not a number about your job. A
2025 survey defines over-reliance as "relying on LLMs beyond their capabilities"[3] and
argues for measurement over impression, which is exactly what the
table above is.

## Implementation details

A team cannot keep one person's notebook. What replaces it is a fixed set of checkable questions,
scored the same way every time and kept as a file: the [evals](/gradient_ascent/techniques/evals/)
topic pointed at a running system rather than at a prompt being drafted. "We have been using it a
while and it seems fine" is the thing the file exists to replace.

Three properties make a team record worth keeping:

- **Broken out by task kind**, the way this site's own eval set splits its questions, so a good
  score on easy lookups cannot stand in for the record on the cases nobody re-checked.
- **Versioned by what produced it**: model id, prompt version, tool set. A rise or a fall is
  only informative if you can name what changed; without that, the record is a mood.
- **Fed automatically where it can be.** A result that carries a citation, the way
  [RAG](/gradient_ascent/techniques/rag/)'s does, can have the mechanical part checked by
  machine: the reviewing page's own checker reports whether a figure the answer states appears in
  the section it cites. That is one row of evidence per answer without a person reading the whole
  thing, and it is a presence check, so it is a floor under the record, not the record.

Sample by hand on top of that, at a rate you write down. The machine check and the human sample
answer different questions, and an automated number rising while nobody has read an output in
six weeks is exactly the state that looks safest and is not.

One design decision belongs here rather than in the record: what the system does when its own
confidence is low. A system that can say "I did not find this" gives a reviewer a signal worth
logging; one that always produces an answer makes every result look identical from outside, and
a record over identical-looking results is much more expensive to keep.

## At each level

- [Conventional software](/gradient_ascent/levels/0/): nothing to calibrate: a rule passes its tests or it
  does not, and it behaves the same way tomorrow.
- [Direct prompting](/gradient_ascent/levels/1/): every reply is independent, so the record is simply a
  tally per task type with nothing else to attribute a change to.
- [Added context](/gradient_ascent/levels/2/): the record has to separate "found the right source and
  read it wrong" from "never found it", because those two have different fixes.
- [Workflows](/gradient_ascent/levels/3/): the steps are fixed, so trust can be tracked per
  step, and one unreliable stage stops dragging down the ones around it.
- [Tool use](/gradient_ascent/levels/4/): the record now needs to cover the action taken, not only
  the sentence produced: a wrong call can be reported in perfectly correct prose.
- [Agent loops](/gradient_ascent/levels/5/): the unit becomes a whole run of variable length, so the
  record tracks outcomes and cost per run rather than accuracy per answer.
- [Teams of Agents](/gradient_ascent/levels/6/): agreement between agents is not evidence. The
  record has to cover the division of labour, since a team can be confidently wrong together.
- [Always-on agents](/gradient_ascent/levels/7/): the record is the only thing standing in for
  someone watching, so it has to be written by the system itself and read by a person on a
  schedule.

## Practices

- Keep the record per task type. "It has been good lately" is not a record.
- Log the checks that came back fine as well as the ones that did not; otherwise the record only
  contains disasters and reads like one.
- Set the sampling rate from the last month's edit rate, and write the number down where someone
  else can see it.
- Re-check after any change to the model, the prompt or the tools, and mark the record with what
  changed.
- Name the tasks you deliberately do not trust, and what evidence would change that. An untested
  assumption in the cautious direction is still an untested assumption.

## Run it

**What to monitor.** Per task type, the edit rate (how often a result is accepted unchanged) next to the
  sampling rate actually being used. Those two numbers moving apart is the whole subject of this
  page, in either direction.

**Cost at volume.** Keeping the record costs roughly the same per checked result at any volume, so
  the sampling rate, not the traffic, sets the bill. What volume changes is the cost of being
  miscalibrated: the same error rate is a nuisance at ten results a day and a recall at ten
  thousand.

**How it fails in production.** The record stops being written before it stops being cited. Months later a
  decision is justified with 'it has been reliable', and the last entry anyone made was before
  two model upgrades and a prompt rewrite.

**What to log.** Every checked result against what was actually correct, tagged with the task type
  and with the model and prompt version that produced it. Untagged accuracy cannot be compared
  across a change, which is the only comparison that matters.

## Try it

1. **Use it.** Pick a task you now let a model do without checking. Write down when you last verified one of its answers and what you found. If you cannot remember, that gap is the distance between your trust and your record, and the next five results are the cheapest rows you will ever add.
2. **Build it.** Take one output type from a system you use or built. Sketch the automatic check for it (what would a machine compare against what) and say what that check would NOT catch. The second half is what the human sample is for.
3. **Either lane.** Name one task you trust a model on and one you deliberately do not. For each, write the evidence the position rests on. If either answer is a feeling rather than a count, that is the one to start a record for.


## Sources

1. [Prompting best practices](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices) — Anthropic (Claude Platform Docs) (accessed 2026-09-19)
2. [Automation-Induced Complacency Potential: Development and Validation of a New Scale](https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2019.00225/full) — Frontiers in Psychology (Merritt et al.), 2019-02-19 (accessed 2026-09-19)
3. [Measuring and mitigating overreliance to build human-compatible AI](https://arxiv.org/abs/2509.08010) — arXiv (Ibrahim et al.), 2025-09-08 (accessed 2026-09-19)


Last reviewed 2026-09-19.
