Reusable instructions that an agent loads when it needs them.
Sourced
Concept at a glance
Load the instructions when the task needs them.
SequenceConceptual illustration
Read the connections in words
Task + skill list → Select a skill: Load the relevant guidance.
Select a skill → Agent uses it: Follow it with available tools.
Key idea
A skill supplies reusable instructions and resources; the agent still has to do the work.
A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.
GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions
Skills: see it in practice.
Reusable task instructions and supporting resources loaded when an agent needs a particular procedure.
What you’ll walk through
Follow a reusable procedure being selected and applied to a specific task. Inspect what the procedure contributes and where current context requires judgment rather than mechanical copying.
The task in this version
Create a release note using our writing procedure.
What you’ll learn to check
Selected skill, loaded resources, draft release note, verification checklist, and outdated-skill handling.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example
Create a release note using our writing procedure.
Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Resources: audience guide, template, verification checklist. Change record: CSV export for empty rows fixed; no new integrations.
What changed: Establish the facts supplied for this version of the task.
WHY THIS MATTERS
What this case assumes
A procedure has a scope, prerequisites, and a version. Instructions inside a skill do not make its outputs correct or authorize unrelated actions.
1 / 6
Apply this to your project
Describe your task to your own model and use Skills as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
Go deeper: practical guidance, failure modes, and implementation
A skill is instructions an agent keeps on the shelf until it needs them, not text it carries into
every turn. Anthropic’s own documentation draws that line directly against a system prompt:
“Unlike prompts (conversation-level instructions for one-off tasks), Skills load on demand, so you
don’t have to repeat the same guidance across conversations”[1]. The mechanism is
progressive disclosure, which Anthropic documents in three stages: a skill’s name and description
load at startup and stay in context; the body of its SKILL.md, the actual instructions, loads
only when the skill is triggered; anything else it bundles (further reference files, scripts)
costs nothing until it is read or run[1].
A skill differs from RAG in what gets fetched (RAG retrieves
facts, a skill supplies a procedure) and from fine-tuning
in where the knowledge lives: fine-tuning bakes a behavior into the weights, while a skill loads
per request and can be edited or removed in seconds.
Choosing which skill to load, if any, is the level-5 decision here, the same shape as choosing a
tool in single agent. Your code runs the lookup,
appends the body, and can cap how many skills load before forcing an answer.
This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.
Optional: inspect the implementation trace
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
Skills
A short description is always in context; the model chooses one and the body loads only then.
Level 5 · Agent loops
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step
The run, step by step
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
STEP 01 / 05Your code chose
The question arrives
"Is the DW-480 drain pump covered under warranty,
and for how long?"
0 tokens · 0 ms
Practical guidance
Skills mostly live in coding agents today: Claude Code’s own skills directory, or a similar folder
any agent that reads an AGENTS.md-style file will pick up. Look for a skills folder in a project,
each entry a short file describing what it does and a longer body of instructions underneath.
Before you add or run one from outside your own team, read the whole body, not just its one-line
description: Anthropic’s own warning is direct, “a malicious Skill can direct Claude to invoke
tools or execute code in ways that don’t match the Skill’s stated purpose,” naming “data
exfiltration” and “unauthorized system access” among the risks[1]. A safe skill next to a
dangerous one can carry an equally short, equally reasonable-sounding description; the difference
only shows up in the body, and in what it asks to be allowed to do.
That last part is a specific field worth reading on its own. Claude Code’s own documentation:
“A skill can grant itself broad tool access, so review the allowed-tools of skills checked into
a repository before you run Claude Code there”[2]. A skill that only reads and
summarizes needs little; one that lists broad file or network access needs the same scrutiny you
would give a code change from someone you don’t know, every time it changes, not just when you
first add it.
To test whether a skill actually loaded rather than the agent answering from general knowledge,
ask it to do the specific task the skill’s description promises, then ask it to name which skill it
used and quote a line straight from that skill’s own instructions, not a summary in its own words.
An answer that cannot quote anything from the file it claims to have read probably never loaded it
at all, and answered from general training instead.
AGENTS.md is the closest thing to a shared convention for the plainer version of this, a project’s
own standing instructions rather than a library of separate skills: “the extra, sometimes detailed
context coding agents need: build steps, tests, and conventions”[3]. If there is only one
procedure you ever need, that file, or a fixed instruction in your prompt, does the job without a
skill library at all.
Implementation details
The example holds three skills in a small registry, each a Skill(name, description, body). Every
description is folded into the system prompt up front: cheap, and always available for the model
to match against. No body is. The model reads the question, decides whether one of the three
descriptions actually fits, and if so calls load_skill with its name; your code looks it up and
appends the body as a new message, and only then has the model actually seen the instructions it
chose.
Every load_skill call, and the decision to stop, are decided_by: "model"; listing the
descriptions up front and loading a chosen body are always decided_by: "code", the same split
single agent’s example makes between what the model
picks and what the program carries out. An unknown skill name is reported back rather than raised,
so a model that guesses a name that does not exist gets a chance to try again instead of crashing
the run. max_steps (3) and max_tokens (1500) are the hard caps; hitting either forces a final
answer that is decided_by: "code".
examples/skills/run.py · lines 66–113
def run(
question: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
registry: dict[str, Skill] = SKILLS,
max_steps: int = MAX_STEPS,
max_tokens: int = MAX_TOKENS,
) -> Answer:
del embedder # skills are loaded from the registry below, not retrieved from the corpus
listing = _skill_list(registry)
tracer.record(kind="code", decided_by="code", title="List skill descriptions (always in context)", detail=listing)
system = (
"Answer the question. These skills are available; each description says what it is for. "
"Call load_skill with a skill's name if one of them applies before answering:\n" + listing
)
messages = [Message(role="system", content=system), Message(role="user", content=question)]
loaded: list[str] = []
tokens_used = 0
for _ in range(max_steps):
completion = model.complete(messages, tools=[LOAD_SKILL_TOOL], max_tokens=250)
tokens_used += completion.tokens_in + completion.tokens_out
if not completion.tool_calls:
record_completion(tracer, decided_by="model", title="Model answers", completion=completion)
return Answer(text=completion.text, citations=loaded)
name = str(completion.tool_calls[0].arguments.get("name", ""))
record_completion(tracer, decided_by="model", title="Model chooses a skill to load", completion=completion, detail=name)
skill = registry.get(name)
if skill is None:
body = f"unknown skill: {name}"
else:
body = skill.body
loaded.append(skill.name)
tracer.record(kind="code", decided_by="code", title="Load the skill body into context", detail=body[:200])
messages.append(Message(role="assistant", content=f"[loaded skill {name}]"))
messages.append(Message(role="user", content=f"Skill '{name}' body:\n{body}"))
if tokens_used >= max_tokens:
reason = f"token budget reached: {tokens_used} >= {max_tokens}"
final = force_final(messages, model, tracer, reason=reason, max_tokens=300)
return Answer(text=final.text, citations=loaded)
final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=300)
return Answer(text=final.text, citations=loaded)
Run it yourself:
examples/skills/README.md · lines 17–17
python -m examples.skills --model stub:scripted
When you do not need this
Try a plain system prompt (prompt engineering)
first if there is only one procedure the agent ever needs; paying to keep even a short
description of it in context buys nothing a fixed instruction does not already give you for free.
Try routing instead if you can tell, from the question
alone and in code, which instructions apply: a classifier that always attaches the same fixed
text for the same category is cheaper and more predictable than asking the model to choose.
Move up to skills once there are several distinct procedures, the right one depends on judgment a
fixed classifier cannot make reliably, and loading every one of them on every request would waste
more context than it is worth.
Failure modes
The wrong skill gets picked
How to notice it
The model matches a description on a surface keyword rather than what the question actually needs, and loads instructions that do not fit the task.
How to test for it
Write two descriptions that share a word but cover different needs (this page's own unit-conversion and warranty-checklist skills are deliberately distinct) and confirm the model's choice tracks the actual task, not just shared vocabulary.
A skill does what its description does not say
How to notice it
Anthropic warns directly that "a malicious Skill can direct Claude to invoke tools or execute code in ways that don't match the Skill's stated purpose."
How to test for it
Read the full body of a skill before trusting its description, especially one from outside your own team; the description is what gets matched against, not what necessarily runs.
A checked-in skill is never actually reviewed
How to notice it
A skill's permissions or instructions change over time the way any file in a repository can, without the same review a code change would get.
How to test for it
Check whether a skill went through the same review as the code around it before it was trusted, not just when it was first added.
A loaded skill goes unused
How to notice it
The model loads a skill's body, spending the tokens progressive disclosure was supposed to save, and then answers without actually following it.
How to test for it
Compare the loaded skill's instructions against what the final answer actually did. A load with no visible effect on the answer is wasted, not just unnecessary.
The cap ships an answer with no skill loaded
How to notice it
A step or token cap is reached before the model ever loaded the skill the task needed, and the forced answer goes out without it.
How to test for it
Force a low cap on a question that needs a skill (this page's own test suite does exactly this) and check whether the returned citations list shows a skill was actually loaded.
Cost and latency
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
2Model calls, one skill loaded then answer
~90Tokens in, description listing
~40Tokens in, one loaded body
~0.7sWall time, one round trip
Compared with always including every skill body (no progressive disclosure)Loading only the chosen skill keeps the unused bodies out of the prompt entirely; with three small skills the saving is modest, but Anthropic's own figures put a real skill body at under 5,000 tokens, so the saving grows with how many skills a project actually has.
How to Evaluate It
skills does not do the site’s own question-answering task (there is no document, no question
about it, and no citation to grade), so it is not part of the shared 60-question set, the same way
embeddings_search and memory are not (see docs/EVALS.md). What would actually be measured is
specific to this task: given a labeled set of questions and the skill each one should trigger,
selection accuracy (did it load the right one, none when none applies, and not an extra one it
never used) and the token cost actually spent against what always-loading every skill would have
cost. scripts/eval_run.py knows skills and refuses to score it, printing that reason; no
runner for the measures above exists yet.
Run it
What to monitor
Skill selection accuracy against a labeled set of tasks, how often a loaded skill's instructions show up in the final answer, and how often the model loads a skill and then ignores it.
Cost at volume
Cost scales with how many requests actually trigger a skill load, not with how many skills exist in the registry: a large registry with rare triggers costs close to nothing per request until one is chosen.
How it fails in production
A skill checked into the project changes without review, and the next session that triggers it inherits whatever it now says or whatever tool access it now grants, with no signal that anything changed.
What to log
The full list of descriptions offered, which skill (if any) was chosen, its complete body at the time it was loaded, and which cap (if any) forced the final answer.
Try it
Use it
Find a project that ships an AGENTS.md, CLAUDE.md or a skills directory. Read one skill's body, not its description: does what it says match what the description promised?
Build it
Run python -m examples.skills --model stub:scripted from the repo root: three descriptions sit in context, the model loads warranty-checklist, and answers from its three steps. Change the name it loads in SCRIPTED (examples/skills/__main__.py) to one that does not exist: the loader reports unknown skill as text, the model answers anyway, and skills loaded reads none.
Either lane
Write two skill descriptions for a task you do regularly, worded so a person skimming them, not just a model, could tell which applies when. If you cannot tell them apart quickly, a model will not either.
Either lane
Write a SKILL.md-shaped body, a numbered procedure rather than prose, for a measurement you would otherwise explain aloud: the ripple on a switching regulator, on the oscilloscope with its 20 MHz bandwidth limit on, never on a multimeter whose AC volts function stops below the frequencies that make the ripple. Could an engineer who never watched you do it follow it? That is the test of a skill.