A record for one person is five columns, takes about twenty seconds a row, and lives wherever
you already keep notes:
| Date |
Task type |
What you checked |
Right? |
What you changed |
| 3/12/2027 |
renewal dates from a contract |
all 4 dates against the clauses |
yes |
nothing |
| 3/12/2027 |
plain-language summary of a policy |
the 2 exclusions it listed |
no |
added the third exclusion |
Task type keeps the record usable later: separate lines even when the same product did both.
What you changed is the check itself: nothing, a word, or the whole thing. After a month,
read down that column for one task type. Mostly “nothing” means you can safely sample instead of
reading every one; anything else means you aren’t ready to stop checking, whatever the product’s
reputation is. Write the number down; it’s next month’s sampling rate, not a feeling you’ll
remember correctly later.
Two rules keep it honest. Log the checks that came back fine, not only the corrections, or the
record reads like a catalog of disasters. And start a fresh page whenever what you’re trusting
changes: a new model version, an edited prompt, a different tool. Anthropic says exactly this
about its own published techniques: where one names a specific model, “treat it as measured on
that model and re-check it against your own evals before applying it to another.”[1] A
record built on last quarter’s version doesn’t transfer just because the product name did.
Skip it for a one-off task you’ll never ask again, or something so low-stakes a wrong answer
costs nothing to fix. Keep it for whatever you catch yourself about to trust from memory instead
of a count.
Be wary of research that sounds like it settles this. A 2019 complacency scale was built on
Mechanical Turk respondents whose “experience with automation was predominantly with relatively
low-stakes and common forms of automation, such as in-car navigation systems,” which its own
authors name as a limitation[2]; those participants weren’t supervising a model at
work, so treat the finding as a reason to keep your own count, not a number about your job. A
2025 survey defines over-reliance as “relying on LLMs beyond their capabilities”[3] and
argues for measurement over impression, which is exactly what the
table above is.