# Skills

_Level 05 · Agent loops · sourced_

Reusable instructions that an agent loads when it needs them.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a reusable procedure being selected and applied to a specific task. Inspect what the procedure contributes and where current context requires judgment rather than mechanical copying.

**Assumptions:** A procedure has a scope, prerequisites, and a version. Instructions inside a skill do not make its outputs correct or authorize unrelated actions.

**Design choices:** Use a skill for recurring know-how; use tools to execute operations. Keep task-specific facts outside the reusable procedure and check whether its assumptions fit.

**Request:** Create a release note using our writing procedure.

**Starting evidence:** Resources: audience guide, template, verification checklist. Change record: CSV export for empty rows fixed; no new integrations.

**Action and control:** Load the reusable procedure and apply it to supplied facts; it grants no publishing authority.

**Stage records (authored, not executed):**

### Input record

Resources: audience guide, template, verification checklist. Change record: CSV export for empty rows fixed; no new integrations.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use a skill for recurring know-how; use tools to execute operations. Keep task-specific facts outside the reusable procedure and check whether its assumptions fit.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Load the reusable procedure and apply it to supplied facts; it grants no publishing authority.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Draft: fixed CSV export for empty rows. Checklist links the claim to the change record; publication remains separate.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Selected skill, loaded resources, draft release note, verification checklist, and outdated-skill handling.

If the result falls short:
If the skill conflicts with current requirements or lacks a necessary step, surface the mismatch and adapt within authority. Do not silently treat an old template as current policy.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to reporting, analysis, design, or repository work. Keep procedures small enough to maintain and make their applicability clear.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Draft: fixed CSV export for empty rows. Checklist links the claim to the change record; publication remains separate.

**Change something — Load an outdated template requiring unsupported claims:** Flag stale instructions rather than inventing an integration to satisfy the template.

**Decision:** Can a skill grant itself permission to publish?

**Answer:** No; authorization remains separate.

**Why:** A skill provides guidance rather than new permissions or guaranteed correctness; stale instructions can conflict with current policy.

**Review criteria:** Selected skill, loaded resources, draft release note, verification checklist, and outdated-skill handling.

**Recovery:** If the skill conflicts with current requirements or lacks a necessary step, surface the mismatch and adapt within authority. Do not silently treat an old template as current policy.

**Adapt it:** Apply this to reporting, analysis, design, or repository work. Keep procedures small enough to maintain and make their applicability clear.

A skill is instructions an agent keeps on the shelf until it needs them, not text it carries into
every turn. Anthropic's own documentation draws that line directly against a system prompt:
"Unlike prompts (conversation-level instructions for one-off tasks), Skills load on demand, so you
don't have to repeat the same guidance across conversations"[1]. The mechanism is
progressive disclosure, which Anthropic documents in three stages: a skill's name and description
load at startup and stay in context; the body of its `SKILL.md`, the actual instructions, loads
only when the skill is triggered; anything else it bundles (further reference files, scripts)
costs nothing until it is read or run[1].

A skill differs from [RAG](/gradient_ascent/techniques/rag/) in what gets fetched (RAG retrieves
facts, a skill supplies a procedure) and from [fine-tuning](/gradient_ascent/techniques/adaptation/)
in where the knowledge lives: fine-tuning bakes a behavior into the weights, while a skill loads
per request and can be edited or removed in seconds.

Choosing which skill to load, if any, is the level-5 decision here, the same shape as choosing a
tool in [single agent](/gradient_ascent/techniques/single-agent/). Your code runs the lookup,
appends the body, and can cap how many skills load before forcing an answer.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 5 · Skills. The same steps are described in the sections below._

## Practical guidance

Skills mostly live in coding agents today: Claude Code's own skills directory, or a similar folder
any agent that reads an AGENTS.md-style file will pick up. Look for a skills folder in a project,
each entry a short file describing what it does and a longer body of instructions underneath.

Before you add or run one from outside your own team, read the whole body, not just its one-line
description: Anthropic's own warning is direct, "a malicious Skill can direct Claude to invoke
tools or execute code in ways that don't match the Skill's stated purpose," naming "data
exfiltration" and "unauthorized system access" among the risks[1]. A safe skill next to a
dangerous one can carry an equally short, equally reasonable-sounding description; the difference
only shows up in the body, and in what it asks to be allowed to do.

That last part is a specific field worth reading on its own. Claude Code's own documentation:
"A skill can grant itself broad tool access, so review the `allowed-tools` of skills checked into
a repository before you run Claude Code there"[2]. A skill that only reads and
summarizes needs little; one that lists broad file or network access needs the same scrutiny you
would give a code change from someone you don't know, every time it changes, not just when you
first add it.

To test whether a skill actually loaded rather than the agent answering from general knowledge,
ask it to do the specific task the skill's description promises, then ask it to name which skill it
used and quote a line straight from that skill's own instructions, not a summary in its own words.
An answer that cannot quote anything from the file it claims to have read probably never loaded it
at all, and answered from general training instead.

AGENTS.md is the closest thing to a shared convention for the plainer version of this, a project's
own standing instructions rather than a library of separate skills: "the extra, sometimes detailed
context coding agents need: build steps, tests, and conventions"[3]. If there is only one
procedure you ever need, that file, or a fixed instruction in your prompt, does the job without a
skill library at all.

## Implementation details

The example holds three skills in a small registry, each a `Skill(name, description, body)`. Every
description is folded into the system prompt up front: cheap, and always available for the model
to match against. No body is. The model reads the question, decides whether one of the three
descriptions actually fits, and if so calls `load_skill` with its name; your code looks it up and
appends the body as a new message, and only then has the model actually seen the instructions it
chose.

Every `load_skill` call, and the decision to stop, are `decided_by: "model"`; listing the
descriptions up front and loading a chosen body are always `decided_by: "code"`, the same split
[single agent](/gradient_ascent/techniques/single-agent/)'s example makes between what the model
picks and what the program carries out. An unknown skill name is reported back rather than raised,
so a model that guesses a name that does not exist gets a chance to try again instead of crashing
the run. `max_steps` (3) and `max_tokens` (1500) are the hard caps; hitting either forces a final
answer that is `decided_by: "code"`.

`examples/skills/run.py` (lines 66-113)

```python
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    registry: dict[str, Skill] = SKILLS,
    max_steps: int = MAX_STEPS,
    max_tokens: int = MAX_TOKENS,
) -> Answer:
    del embedder  # skills are loaded from the registry below, not retrieved from the corpus
    listing = _skill_list(registry)
    tracer.record(kind="code", decided_by="code", title="List skill descriptions (always in context)", detail=listing)
    system = (
        "Answer the question. These skills are available; each description says what it is for. "
        "Call load_skill with a skill's name if one of them applies before answering:\n" + listing
    )
    messages = [Message(role="system", content=system), Message(role="user", content=question)]

    loaded: list[str] = []
    tokens_used = 0
    for _ in range(max_steps):
        completion = model.complete(messages, tools=[LOAD_SKILL_TOOL], max_tokens=250)
        tokens_used += completion.tokens_in + completion.tokens_out

        if not completion.tool_calls:
            record_completion(tracer, decided_by="model", title="Model answers", completion=completion)
            return Answer(text=completion.text, citations=loaded)

        name = str(completion.tool_calls[0].arguments.get("name", ""))
        record_completion(tracer, decided_by="model", title="Model chooses a skill to load", completion=completion, detail=name)
        skill = registry.get(name)
        if skill is None:
            body = f"unknown skill: {name}"
        else:
            body = skill.body
            loaded.append(skill.name)
        tracer.record(kind="code", decided_by="code", title="Load the skill body into context", detail=body[:200])
        messages.append(Message(role="assistant", content=f"[loaded skill {name}]"))
        messages.append(Message(role="user", content=f"Skill '{name}' body:\n{body}"))

        if tokens_used >= max_tokens:
            reason = f"token budget reached: {tokens_used} >= {max_tokens}"
            final = force_final(messages, model, tracer, reason=reason, max_tokens=300)
            return Answer(text=final.text, citations=loaded)

    final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=300)
    return Answer(text=final.text, citations=loaded)
```

Run it yourself:

`examples/skills/README.md` (lines 17-17)

```text
python -m examples.skills --model stub:scripted
```

## When you do not need this

Try a plain system prompt ([prompt engineering](/gradient_ascent/techniques/prompt-engineering/))
first if there is only one procedure the agent ever needs; paying to keep even a short
description of it in context buys nothing a fixed instruction does not already give you for free.

Try [routing](/gradient_ascent/techniques/routing/) instead if you can tell, from the question
alone and in code, which instructions apply: a classifier that always attaches the same fixed
text for the same category is cheaper and more predictable than asking the model to choose.

Move up to skills once there are several distinct procedures, the right one depends on judgment a
fixed classifier cannot make reliably, and loading every one of them on every request would waste
more context than it is worth.

## Failure modes

### The wrong skill gets picked

- **How to notice it:** The model matches a description on a surface keyword rather than what the question actually needs, and loads instructions that do not fit the task.
- **How to test for it:** Write two descriptions that share a word but cover different needs (this page's own unit-conversion and warranty-checklist skills are deliberately distinct) and confirm the model's choice tracks the actual task, not just shared vocabulary.

### A skill does what its description does not say

- **How to notice it:** Anthropic warns directly that "a malicious Skill can direct Claude to invoke tools or execute code in ways that don't match the Skill's stated purpose."
- **How to test for it:** Read the full body of a skill before trusting its description, especially one from outside your own team; the description is what gets matched against, not what necessarily runs.

### A checked-in skill is never actually reviewed

- **How to notice it:** A skill's permissions or instructions change over time the way any file in a repository can, without the same review a code change would get.
- **How to test for it:** Check whether a skill went through the same review as the code around it before it was trusted, not just when it was first added.

### A loaded skill goes unused

- **How to notice it:** The model loads a skill's body, spending the tokens progressive disclosure was supposed to save, and then answers without actually following it.
- **How to test for it:** Compare the loaded skill's instructions against what the final answer actually did. A load with no visible effect on the answer is wasted, not just unnecessary.

### The cap ships an answer with no skill loaded

- **How to notice it:** A step or token cap is reached before the model ever loaded the skill the task needed, and the forced answer goes out without it.
- **How to test for it:** Force a low cap on a question that needs a skill (this page's own test suite does exactly this) and check whether the returned citations list shows a skill was actually loaded.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one skill loaded then answer:** 2
- **Tokens in, description listing:** ~90
- **Tokens in, one loaded body:** ~40
- **Wall time, one round trip:** ~0.7s

**Compared with always including every skill body (no progressive disclosure).** Loading only the chosen skill keeps the unused bodies out of the prompt entirely; with three small skills the saving is modest, but Anthropic's own figures put a real skill body at under 5,000 tokens, so the saving grows with how many skills a project actually has.

## How to Evaluate It

`skills` does not do the site's own question-answering task (there is no document, no question
about it, and no citation to grade), so it is not part of the shared 60-question set, the same way
`embeddings_search` and `memory` are not (see `docs/EVALS.md`). What would actually be measured is
specific to this task: given a labeled set of questions and the skill each one should trigger,
selection accuracy (did it load the right one, none when none applies, and not an extra one it
never used) and the token cost actually spent against what always-loading every skill would have
cost. `scripts/eval_run.py` knows `skills` and refuses to score it, printing that reason; no
runner for the measures above exists yet.

## Run it

**What to monitor.** Skill selection accuracy against a labeled set of tasks, how often a loaded skill's instructions show up in the final answer, and how often the model loads a skill and then ignores it.

**Cost at volume.** Cost scales with how many requests actually trigger a skill load, not with how many skills exist in the registry: a large registry with rare triggers costs close to nothing per request until one is chosen.

**How it fails in production.** A skill checked into the project changes without review, and the next session that triggers it inherits whatever it now says or whatever tool access it now grants, with no signal that anything changed.

**What to log.** The full list of descriptions offered, which skill (if any) was chosen, its complete body at the time it was loaded, and which cap (if any) forced the final answer.

## Try it

1. **Use it.** Find a project that ships an AGENTS.md, CLAUDE.md or a skills directory. Read one skill's body, not its description: does what it says match what the description promised?
2. **Build it.** Run python -m examples.skills --model stub:scripted from the repo root: three descriptions sit in context, the model loads warranty-checklist, and answers from its three steps. Change the name it loads in SCRIPTED (examples/skills/__main__.py) to one that does not exist: the loader reports unknown skill as text, the model answers anyway, and skills loaded reads none.
3. **Either lane.** Write two skill descriptions for a task you do regularly, worded so a person skimming them, not just a model, could tell which applies when. If you cannot tell them apart quickly, a model will not either.
4. **Either lane.** Write a SKILL.md-shaped body, a numbered procedure rather than prose, for a measurement you would otherwise explain aloud: the ripple on a switching regulator, on the oscilloscope with its 20 MHz bandwidth limit on, never on a multimeter whose AC volts function stops below the frequencies that make the ripple. Could an engineer who never watched you do it follow it? That is the test of a skill.


## Sources

1. [Agent Skills](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview) — Anthropic (accessed 2026-09-19)
2. [Extend Claude with skills](https://code.claude.com/docs/en/skills) — Anthropic (Claude Code documentation) (accessed 2026-09-19)
3. [AGENTS.md](https://agents.md/) — agents.md (stewarded by the Agentic AI Foundation, Linux Foundation) (accessed 2026-09-19)


Last reviewed 2026-09-19.
