Topics at every level

Working with a model

How to brief a model, review its work and decide what to hand over.

Sourced

Concept at a glance

Brief well, inspect the work, own the decision.

SequenceConceptual illustration
Brief well, inspect the work, own the decision.Brief leads to Model does work. Model does work leads to Review + decide. The quality of working with a model depends on what you ask and what you verify.BriefGoal, context, constraintsModel does workWithin the scope you gaveReview + decideCheck before relying on itBrief well, inspect the work, own the decision.Brief leads to Model does work. Model does work leads to Review + decide. The quality of working with a model depends on what you ask and what you verify.BriefGoal, context, constraintsModel does workWithin the scope you gaveReview + decideCheck before relying on it
Read the connections in words
  • Brief → Model does work: Within the scope you gave.
  • Model does work → Review + decide: Check before relying on it.
Key idea

The quality of working with a model depends on what you ask and what you verify.

A focused everyday life example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Working with a model: see it in practice.

Practices for specifying, delegating, reviewing, and evaluating work done with a model.

What you’ll walk through

Follow a person shaping a task, inspecting a draft, and deciding what to accept or revise. The useful skill is directing attention toward the parts that need human judgment.

The task in this version

Help me prepare a workshop plan I can responsibly approve.

What you’ll learn to check

Task brief, delegation boundary, draft, independent checks, and a recorded acceptance decision.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Everyday lifeAn authored case with its own evidence, changed condition, and decision.
The task in this example

Help me prepare a workshop plan I can responsibly approve.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
20 attendees, $300 budget, accessible venue. Assistant drafts; organizer controls bookings.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Fluent output can hide omissions, and the user may not initially know every requirement. The task can become clearer through iteration.

1 / 6
In this topic

4 pages under working with a model

Each one goes further into a part of this page than this page does.

Briefing: saying what you want

Sourced

Saying what you want clearly enough that the model does not have to guess.

Reviewing work you did not do

Sourced

Checking work you did not do yourself before it goes anywhere.

Deciding what to hand over

Sourced

Deciding which parts of a task to hand to a model and which to keep.

Calibrating trust

Sourced

Learning, from results over time, how much to rely on a model without checking.

Apply this to your project

Describe your task to your own model and use Working with a model as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Operator craft is the human half of the manual: the skill of working well with a model, which grows more demanding, not less, as a system climbs the ladder. At level 1 it is mostly one skill, saying clearly what you want. By level 7 it is several: judging work you did not do yourself, deciding what to hand over and what to keep, and knowing, from evidence rather than a good first impression, how much to trust a system that acts without asking each time. This page is the overview; four pages carry the depth: briefing, reviewing, delegating, and calibrating trust.

None of this is specific to a coder. A domain professional who has never written a line of code briefs, reviews, delegates and trusts a model exactly the way a builder does, on the same four skills, just without the code.

This page is sourced, not measured: the habits below are checked against primary sources, and none of them has been tried on a scored task here.

Practical guidance

Four things to do this week, one per skill, each handed to its own page for the full depth.

Briefing: before your next real request, write the goal, the constraint and what “done” looks like on three separate lines instead of folding them into one paragraph. Anthropic’s own guidance states the underlying point plainly: “Claude responds well to clear, explicit instructions”[1]. How much detail a model needs also depends on which one you’re briefing: OpenAI’s own guidance sorts its models into two kinds for exactly that reason, one that works out the details from a goal on its own and one that needs them spelled out[2]. Read a result that missed the mark against those three lines before rewriting it louder; the briefing page has the full six-part version and a worked example.

Reviewing: pick one thing you routinely accept from a model without checking, and check it once this week against something independent, a citation that resolves, a number you can recompute. You’ll know it worked when you can point to what you verified, not just say the answer “seemed right”; if you can’t find anything in the output to check it against at all, that’s the failure, not your diligence. Anthropic’s own guidance for agent builders treats human checkpoints as a normal part of design, not a fallback, and even for coding, where “Code solutions are verifiable through automated tests,” says human review remains crucial[3]; the reviewing page has the rest.

Delegating: before handing something over, ask what a wrong answer would cost and how you’d notice, then hand over only the part that scores well on both; the delegating page turns those two questions, plus two more, into a short table you can score any task against.

Trust: a string of correct answers is exactly when to keep checking, not stop. A scale-development paper on automation complacency found that monitoring a system “at a frequency that is suboptimal or below a normative rate” leads to performance failures[4]; calibrating trust covers what a record of your own corrections over time should look like, and how to read it.

None of this is worth doing for a single, low-stakes request you wouldn’t mind redoing by hand yourself.

Implementation details

A builder shapes operator craft through the interface, whether or not any of this is written down as a rule. A free-text chat box makes the brief and the request the same blob of prose, so nothing stops the constraint or the definition of done from getting silently dropped on a rewrite. A task input with separate fields (what to do, what not to do, what finished looks like) costs a little more to build and makes the brief a thing the reader can check against later, not just a paragraph they typed and moved past.

Reviewing needs something to review. A system that shows only a final answer gives a reviewer nothing to check but the answer’s surface plausibility; one that shows its sources, and ideally the steps that produced the answer, lets a reviewer check the parts that can be verified independently, the way this site’s own trace player exposes each step of a run instead of only its last one. Anthropic’s guidance for agent builders treats the interface between a model and its tools as real design work rather than plumbing: one rule of thumb it gives is to think about how much effort goes into human-computer interfaces and “plan to invest just as much effort in creating good agent-computer interfaces”[3]. The surface a reviewer reads deserves the same budget.

Approval points are where operator craft becomes a piece of code, not just a habit. An interface that lets a person approve or reject one specific action before it runs (rather than trusting an instruction in a system prompt to hold) is the same shape as the permission check on the safety, privacy and governance page: a check outside the model, that runs whether or not the model would have gotten it right on its own. A builder designing for delegation decides, ahead of time and in code, which actions get that checkpoint and which do not; leaving that decision to be made informally, in the moment, is how it stops getting made at all once volume goes up.

No runnable example accompanies this page. A task-input form, a review surface, and an approval checkpoint are interface and process decisions, not an algorithm with a fixed answer to test against: the honest Build it lane here is what to design for, not code to run.

When you do not need this

Skip building any of this out formally (a task-input form with separate fields, a visible review surface, a coded approval checkpoint) for a single, low-stakes request you would not mind redoing by hand. Writing the constraint and the “done” condition on their own line, the way the Practices below ask, is already most of the value, and it costs nothing to try before building anything around it.

There is nothing to brief, review, delegate or trust before a model is actually in the loop; level 0, no model at all does not raise any of these questions.

Build the formal version once a wrong answer costs enough, or gets reviewed by someone other than the person who asked, that “I would have caught it” stops being a good enough plan.

Failure modes

A missing constraint is invisible until it is violated

How to notice it
A request typed as one paragraph drops a constraint on a rewrite with nothing showing it went missing, and the answer that comes back looks fine until the dropped constraint turns out to matter.
How to test for it
Compare a request's current wording against its original list of what, what not, and what done looks like; a constraint no longer present anywhere is this failure, not a model that ignored it.

Automation complacency

How to notice it
After a run of correct answers, checking starts to feel like wasted effort, and the one wrong answer that actually matters goes through with less scrutiny than the ones before it, not more.
How to test for it
Track how often a problem is caught after the fact instead of during review, over time; a rising after-the-fact rate with no change in review effort is this failure, already underway.

Over-specifying a model that could have worked it out

How to notice it
A brief spells out steps the model would have chosen correctly on its own, and the extra constraints leave it less room to handle a case the brief's author did not think to cover.
How to test for it
Compare the outcome of a detailed, step-by-step brief against a shorter one stating only the goal and the constraints, on a task the model has handled well before.

Under-specifying a model that needed the detail spelled out

How to notice it
A brief states only the goal, and the model fills the gap with a plausible-sounding assumption instead of asking, on a task that actually needed a constraint spelled out.
How to test for it
Read the output for an assumption nowhere in the brief; a model that filled a real gap silently, rather than flagging it, is this failure regardless of whether the assumption happened to be right.

Nothing in the interface gives a reviewer something to check

How to notice it
A tool shows only a final answer, so a reviewer can judge no more than whether it sounds plausible, and a wrong answer that reads fluently passes review the same as a right one would.
How to test for it
Try to verify one specific claim in the output independently of the tool itself: a citation, a recomputed number. If the interface gives you nothing to check it against, that is the failure, not the reviewer's diligence.

At each level

  • Conventional software: there is nothing to brief, review, delegate or trust yet; this topic’s skill only starts once a model is actually in the loop.
  • Direct prompting: briefing is nearly the whole skill, since a single prompt is the entire interface between what you want and what you get back.
  • Added context: reviewing has to include what the model was given, not only what it said: RAG’s own failure modes show a wrong answer built from the right sources is a different problem than one built from missing ones, and telling the two apart is its own skill.
  • Workflows: a fixed pipeline hands you a defined checkpoint, the way human approval’s own pause-and-resume shape does, so reviewing can happen at a point along the way instead of only at the end.
  • Tool use: delegating a specific action, not just an answer, becomes a real decision (function calling is exactly that decision in code), and trusting a model to describe what it would do is not the same judgment as trusting it to actually do it.
  • Agent loops: the model chooses its own steps, so calibrating trust (how much to check, and how often, instead of checking every single step a single agent takes) replaces reviewing each one by hand.
  • Teams of Agents: reviewing shifts toward reviewing a division of labor, not one output: whether a lead’s split of the task was sensible is a separate question from whether each part came back right.
  • Always-on agents: delegating reaches its hardest form: deciding in advance what an agent may do without asking and what it must always ask about, the three-way policy always-on assistants’ own example builds, since there is no longer a moment where a person is watching in real time to decide.

Practices

  • Write the constraint and the “done” condition down separately from the request itself, not folded into one paragraph that’s easy to reread without noticing a piece went missing.
  • When reviewing work you didn’t do, check what you can verify independently first (a citation that resolves, a number that recomputes) before trusting the parts you can’t check that way.
  • Decide what to hand over by what a wrong answer costs and how fast you’d notice it: the same question level 0, no model at all asks about choosing a level at all, applied here to one task instead.
  • Calibrate trust from a running record of corrections over time, not from how confident the most recent answer sounded.
  • Treat a string of correct answers as a reason to keep sample-checking, not a reason to stop: that is exactly where the research on automation complacency says monitoring erodes.

Run it

What to monitor

How often you catch a problem after the fact instead of before, on work you were supposed to be reviewing as it happened. A rising after-the-fact rate is the signal to review more closely, not a sign that less review is now safe.

Cost at volume

Review time should not fall to zero as trust grows; it should shift from checking everything to checking a sample plus anything that crosses a fixed bar (an irreversible action, a number that gets repeated elsewhere), the same shape a human-in-the-loop checkpoint uses.

How it fails in production

Automation complacency: after enough correct answers in a row, checking starts to feel like wasted effort, and the one wrong answer that actually mattered goes through unchecked.

What to log

What you changed or rejected in anything you reviewed, and why. A record of your own corrections, not a memory of how confident things felt, is what calibrated trust is actually built from.

Try it

  1. Use it

    Before your next real request to a model, write the constraint and the "done" condition on their own line, separate from the request. Did writing them down change what you actually asked for?

  2. Build it

    Look at a tool you use or have built that shows a model's reasoning or sources before its final answer. Find one place in it where a wrong intermediate step would be invisible to a reviewer, and write down what would have to change to surface it.

  3. Either lane

    Pick one thing you routinely accept from a model without checking. Write down what checking it would actually cost you in time, and what a wrong one slipping through would cost. Decide, on paper, whether that's the trade you actually want.

Optional: products, tools, and models

Concrete examples

In practice

Use a model for a draft, own the outcome

Give it a clear brief, inspect the result against evidence, and decide what is ready to use.

An illustrative task example. No verified product or tool is currently listed for this concept.

Out there

Named products, tools and models

No product, tool or model is registered against this page yet. The names index lists every one the site does name, and which technique each belongs to.

Open the names index →

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Prompting best practices · Anthropic (Claude Platform Docs) (accessed 09/19/2026)
  2. Prompt engineering · OpenAI (API documentation) (accessed 09/19/2026)
  3. Building Effective AI Agents · Anthropic (accessed 09/19/2026)
  4. Automation-Induced Complacency Potential: Development and Validation of a New Scale · Frontiers in Psychology, 02/19/2019 (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page