Primary sources
- Prompting best practices · Anthropic (Claude Platform Docs) (accessed 09/19/2026)
- LLM01:2025 Prompt Injection · OWASP Gen AI Security Project (accessed 09/19/2026)
- Building Effective AI Agents · Anthropic (accessed 09/19/2026)
Deciding which parts of a task to hand to a model and which to keep.
Sourced
Concept at a glance
Scope and review should match the consequence of a mistake.
Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.
Choosing which work a model may perform and which decisions or actions a person retains.
Follow a task being divided between an assistant and a person. Inspect what can proceed independently and what should return as a proposal or question.
Help organize the workshop, but let me control commitments.
Delegation matrix, permitted drafts, withheld transaction, and a change-of-scope approval case.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Help organize the workshop, but let me control commitments.
Authored case. Select any record below; nothing is sent to a model.What changed: Establish the facts supplied for this version of the task.
Capability and authority are separate. The assistant may be able to perform an action that the user only asked it to prepare.
Describe your task to your own model and use Deciding what to hand over as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
Delegating is deciding which parts of a task to hand to a model and which to keep. It comes before briefing, which assumes the handoff has already been settled, and it is a decision about the task rather than about the model: a system fully capable of drafting a refund email can still be the wrong thing to let send one unattended, if a wrong send is expensive and hard to undo.
Four questions do most of the work. What would a wrong answer cost? How easily could you check the result? How reversible is the action once taken? Does the task need context only you have? None of the four asks how good the model is.
Anthropic’s prompting guide draws the same line. For teams who want a model to confirm before risky actions it publishes a sample prompt: text you add to your own, not a description of default behavior: “Consider the reversibility and potential impact of your actions. You are encouraged to take local, reversible actions like editing files or running tests, but for actions that are hard to reverse, affect shared systems, or could be destructive, ask the user before proceeding.”[1]
This page is sourced, not measured: the advice below is checked against primary sources, but no result file exists for any of it, so no number here is one this site took.
Score the task, not the model. Each row is a question about the work in front of you, answered before you hand anything over.
| Question | Hand it over | Keep it, or check before it acts |
|---|---|---|
| What does a wrong answer cost? | A rough draft, a first pass, something you were going to redo anyway | Money moves, a claim goes out, a diagnosis or an order is recorded |
| How checkable is the result? | You can verify it in less time than doing it took: a date, a total, a citation | A judgment call, or a summary of more material than you will reread |
| How reversible is the action? | A draft, a file, a search, a suggestion | A payment, a deletion, a message already sent, a filing already made |
| Whose context does it need? | Everything relevant is in the material you can give it | It turns on a relationship, an unwritten exception, or something said in a room |
Read the row that scores worst, not the average. A task on the right of even one row wants a person between the model and the action, not more trust but a checkpoint; a task on the left of all four is reasonable to hand over unattended.
You’ll know the score was right when a handed-over task keeps coming back cheap to check and cheap to redo. If one starts costing more to verify than it saved, that’s not the model getting worse; it’s the row you scored wrong, and the fix is re-scoring, not adding a step to catch the surprise afterward.
Three things this site recommends never running unattended, regardless of score: anything that moves money or creates a legal obligation on your behalf, anything that communicates in your name outside your own team, and anything that deletes or overwrites the only copy of something. Each fails the same two rows: irreversible, and unverifiable after the fact.
For everything else, widen deliberately: hand over the cheapest, most reversible slice first, watch what actually goes wrong for a few weeks, and widen only past what held up. Skip the framework for a single request you’d happily redo yourself; it earns its cost once a kind of task is going to repeat.
A builder makes the decision durable by encoding it as a permission rather than a habit: a tool the model may call freely, a tool that needs a person’s approval first, and a tool the system never offers at all because nothing in the task needs it. The third category is the one usually skipped. OWASP’s guidance for prompt injection puts both halves plainly: “Implement human-in-the-loop controls for privileged operations to prevent unauthorized actions,” and “Restrict the model’s access privileges to the minimum necessary for its intended operations.”[2] An action the system was never given is one no instruction can talk it into.
The check belongs in code, running whether or not the model would have drawn the boundary correctly by itself. The safety page’s example is one version: a permission check that refuses a refund call unless the customer’s own message independently names the same amount, no matter what a retrieved note tried to talk the model into. The model may propose; only a call the person’s own words support runs.
Deciding what to hand over also decides who else gets to hold it. Anything sitting between you and the model (a workflow automation service, a browser extension, an agent framework’s hosted tracing) receives the same material the model does and holds it under its own terms, not the model maker’s. So the thing to check before routing a task through one is the whole route the material travels, not only the retention page of the company whose model answers. That check is on the safety, privacy and governance page.
Start narrow by default. Anthropic’s own advice to agent builders is “finding the simplest solution possible, and only increasing complexity when needed”[3]: the Use it lane’s “widen deliberately,” stated as a design default rather than a habit somebody has to remember.
Delegating to another agent is still delegating, and the four questions apply to the split as well as to the original task. Anthropic documents over-delegation as a behavior to prompt against rather than a hypothetical: under the heading “Watch for overuse”, its prompting guide says “Claude Opus 5 also delegates to subagents more readily than prior models”, and points to its own sample prompt for damping that down[1]. That is Anthropic describing its own model, and it is worth knowing before reading a run: a split you did not choose is still a delegation, and lead agent and workers is where it gets designed on purpose.
The distance between what a system is permitted to do and what it actually does. Permissions granted and never exercised are the ones to remove; actions attempted and refused are the ones to read, since each is either a boundary working or a boundary in the wrong place.
A permission boundary costs the same whatever the volume: a disallowed action is disallowed whether attempted once or ten thousand times. What grows with volume is the cost of one drawn too loosely, because the mistake now repeats at the rate of the traffic.
The boundary was drawn for a system's first, narrow job and never re-scored as the job widened. Nothing changed in the permissions; what changed is that the actions behind them stopped being cheap and reversible.
Every action taken without asking, the permission that allowed it, and who set that permission and when. The last part is what makes the boundary reviewable by someone other than the person who drew it.
Take a task you already hand to a model and score it on the four questions: cost of a wrong answer, checkability, reversibility, context only you have. Does how closely you actually watch it match the row that scored worst?
Find a tool or agent you use that has an auto-approve or autopilot setting. Turn it off for one session and write down every action it would otherwise have taken without asking. Score each on the four questions; the ones that fail a row are the ones to keep asking about.
Pick one task you have never delegated at all. Score it on the four questions and see whether not delegating is the answer the table gives, or just the default you never revisited.
1 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.
Open-source coding agent
Maker’s documentation Checked 09/18/2026Ask the model to draft a purchasing comparison while a person checks the evidence and makes the purchase decision.
Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page