Level 05 · Agent loops

Skills

Reusable instructions that an agent loads when it needs them.

Sourced

Concept at a glance

Load the instructions when the task needs them.

SequenceConceptual illustration
Load the instructions when the task needs them.Task + skill list leads to Select a skill. Select a skill leads to Agent uses it. A skill supplies reusable instructions and resources; the agent still has to do the work.Task + skill listRead short descriptionsSelect a skillLoad the relevant guidanceAgent uses itFollow it with availabletoolsLoad the instructions when the task needs them.Task + skill list leads to Select a skill. Select a skill leads to Agent uses it. A skill supplies reusable instructions and resources; the agent still has to do the work.Task + skill listRead short descriptionsSelect a skillLoad the relevant guidanceAgent uses itFollow it with availabletools
Read the connections in words
  • Task + skill list → Select a skill: Load the relevant guidance.
  • Select a skill → Agent uses it: Follow it with available tools.
Key idea

A skill supplies reusable instructions and resources; the agent still has to do the work.

A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Skills: see it in practice.

Reusable task instructions and supporting resources loaded when an agent needs a particular procedure.

What you’ll walk through

Follow a reusable procedure being selected and applied to a specific task. Inspect what the procedure contributes and where current context requires judgment rather than mechanical copying.

The task in this version

Create a release note using our writing procedure.

What you’ll learn to check

Selected skill, loaded resources, draft release note, verification checklist, and outdated-skill handling.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Create a release note using our writing procedure.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Resources: audience guide, template, verification checklist. Change record: CSV export for empty rows fixed; no new integrations.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

A procedure has a scope, prerequisites, and a version. Instructions inside a skill do not make its outputs correct or authorize unrelated actions.

1 / 6

Apply this to your project

Describe your task to your own model and use Skills as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

A skill is instructions an agent keeps on the shelf until it needs them, not text it carries into every turn. Anthropic’s own documentation draws that line directly against a system prompt: “Unlike prompts (conversation-level instructions for one-off tasks), Skills load on demand, so you don’t have to repeat the same guidance across conversations”[1]. The mechanism is progressive disclosure, which Anthropic documents in three stages: a skill’s name and description load at startup and stay in context; the body of its SKILL.md, the actual instructions, loads only when the skill is triggered; anything else it bundles (further reference files, scripts) costs nothing until it is read or run[1].

A skill differs from RAG in what gets fetched (RAG retrieves facts, a skill supplies a procedure) and from fine-tuning in where the knowledge lives: fine-tuning bakes a behavior into the weights, while a skill loads per request and can be edited or removed in seconds.

Choosing which skill to load, if any, is the level-5 decision here, the same shape as choosing a tool in single agent. Your code runs the lookup, appends the body, and can cap how many skills load before forcing an answer.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Skills

A short description is always in context; the model chooses one and the body loads only then.

Level 5 · Agent loops
QuestionQuestionlist descriptions onlylistdescriptions onlyMODELpicks the next actionpicks thenext actionTOOLload_skill(name)load_skill(name)AnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 05Your code chose

The question arrives

"Is the DW-480 drain pump covered under warranty,
and for how long?"
0 tokens · 0 ms

Practical guidance

Skills mostly live in coding agents today: Claude Code’s own skills directory, or a similar folder any agent that reads an AGENTS.md-style file will pick up. Look for a skills folder in a project, each entry a short file describing what it does and a longer body of instructions underneath.

Before you add or run one from outside your own team, read the whole body, not just its one-line description: Anthropic’s own warning is direct, “a malicious Skill can direct Claude to invoke tools or execute code in ways that don’t match the Skill’s stated purpose,” naming “data exfiltration” and “unauthorized system access” among the risks[1]. A safe skill next to a dangerous one can carry an equally short, equally reasonable-sounding description; the difference only shows up in the body, and in what it asks to be allowed to do.

That last part is a specific field worth reading on its own. Claude Code’s own documentation: “A skill can grant itself broad tool access, so review the allowed-tools of skills checked into a repository before you run Claude Code there”[2]. A skill that only reads and summarizes needs little; one that lists broad file or network access needs the same scrutiny you would give a code change from someone you don’t know, every time it changes, not just when you first add it.

To test whether a skill actually loaded rather than the agent answering from general knowledge, ask it to do the specific task the skill’s description promises, then ask it to name which skill it used and quote a line straight from that skill’s own instructions, not a summary in its own words. An answer that cannot quote anything from the file it claims to have read probably never loaded it at all, and answered from general training instead.

AGENTS.md is the closest thing to a shared convention for the plainer version of this, a project’s own standing instructions rather than a library of separate skills: “the extra, sometimes detailed context coding agents need: build steps, tests, and conventions”[3]. If there is only one procedure you ever need, that file, or a fixed instruction in your prompt, does the job without a skill library at all.

Implementation details

The example holds three skills in a small registry, each a Skill(name, description, body). Every description is folded into the system prompt up front: cheap, and always available for the model to match against. No body is. The model reads the question, decides whether one of the three descriptions actually fits, and if so calls load_skill with its name; your code looks it up and appends the body as a new message, and only then has the model actually seen the instructions it chose.

Every load_skill call, and the decision to stop, are decided_by: "model"; listing the descriptions up front and loading a chosen body are always decided_by: "code", the same split single agent’s example makes between what the model picks and what the program carries out. An unknown skill name is reported back rather than raised, so a model that guesses a name that does not exist gets a chance to try again instead of crashing the run. max_steps (3) and max_tokens (1500) are the hard caps; hitting either forces a final answer that is decided_by: "code".

examples/skills/run.py · lines 66–113
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    registry: dict[str, Skill] = SKILLS,
    max_steps: int = MAX_STEPS,
    max_tokens: int = MAX_TOKENS,
) -> Answer:
    del embedder  # skills are loaded from the registry below, not retrieved from the corpus
    listing = _skill_list(registry)
    tracer.record(kind="code", decided_by="code", title="List skill descriptions (always in context)", detail=listing)
    system = (
        "Answer the question. These skills are available; each description says what it is for. "
        "Call load_skill with a skill's name if one of them applies before answering:\n" + listing
    )
    messages = [Message(role="system", content=system), Message(role="user", content=question)]

    loaded: list[str] = []
    tokens_used = 0
    for _ in range(max_steps):
        completion = model.complete(messages, tools=[LOAD_SKILL_TOOL], max_tokens=250)
        tokens_used += completion.tokens_in + completion.tokens_out

        if not completion.tool_calls:
            record_completion(tracer, decided_by="model", title="Model answers", completion=completion)
            return Answer(text=completion.text, citations=loaded)

        name = str(completion.tool_calls[0].arguments.get("name", ""))
        record_completion(tracer, decided_by="model", title="Model chooses a skill to load", completion=completion, detail=name)
        skill = registry.get(name)
        if skill is None:
            body = f"unknown skill: {name}"
        else:
            body = skill.body
            loaded.append(skill.name)
        tracer.record(kind="code", decided_by="code", title="Load the skill body into context", detail=body[:200])
        messages.append(Message(role="assistant", content=f"[loaded skill {name}]"))
        messages.append(Message(role="user", content=f"Skill '{name}' body:\n{body}"))

        if tokens_used >= max_tokens:
            reason = f"token budget reached: {tokens_used} >= {max_tokens}"
            final = force_final(messages, model, tracer, reason=reason, max_tokens=300)
            return Answer(text=final.text, citations=loaded)

    final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=300)
    return Answer(text=final.text, citations=loaded)

Run it yourself:

examples/skills/README.md · lines 17–17
python -m examples.skills --model stub:scripted
When you do not need this

Try a plain system prompt (prompt engineering) first if there is only one procedure the agent ever needs; paying to keep even a short description of it in context buys nothing a fixed instruction does not already give you for free.

Try routing instead if you can tell, from the question alone and in code, which instructions apply: a classifier that always attaches the same fixed text for the same category is cheaper and more predictable than asking the model to choose.

Move up to skills once there are several distinct procedures, the right one depends on judgment a fixed classifier cannot make reliably, and loading every one of them on every request would waste more context than it is worth.

Failure modes

The wrong skill gets picked

How to notice it
The model matches a description on a surface keyword rather than what the question actually needs, and loads instructions that do not fit the task.
How to test for it
Write two descriptions that share a word but cover different needs (this page's own unit-conversion and warranty-checklist skills are deliberately distinct) and confirm the model's choice tracks the actual task, not just shared vocabulary.

A skill does what its description does not say

How to notice it
Anthropic warns directly that "a malicious Skill can direct Claude to invoke tools or execute code in ways that don't match the Skill's stated purpose."
How to test for it
Read the full body of a skill before trusting its description, especially one from outside your own team; the description is what gets matched against, not what necessarily runs.

A checked-in skill is never actually reviewed

How to notice it
A skill's permissions or instructions change over time the way any file in a repository can, without the same review a code change would get.
How to test for it
Check whether a skill went through the same review as the code around it before it was trusted, not just when it was first added.

A loaded skill goes unused

How to notice it
The model loads a skill's body, spending the tokens progressive disclosure was supposed to save, and then answers without actually following it.
How to test for it
Compare the loaded skill's instructions against what the final answer actually did. A load with no visible effect on the answer is wasted, not just unnecessary.

The cap ships an answer with no skill loaded

How to notice it
A step or token cap is reached before the model ever loaded the skill the task needed, and the forced answer goes out without it.
How to test for it
Force a low cap on a question that needs a skill (this page's own test suite does exactly this) and check whether the returned citations list shows a skill was actually loaded.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

2Model calls, one skill loaded then answer
~90Tokens in, description listing
~40Tokens in, one loaded body
~0.7sWall time, one round trip
Compared with always including every skill body (no progressive disclosure)Loading only the chosen skill keeps the unused bodies out of the prompt entirely; with three small skills the saving is modest, but Anthropic's own figures put a real skill body at under 5,000 tokens, so the saving grows with how many skills a project actually has.

How to Evaluate It

skills does not do the site’s own question-answering task (there is no document, no question about it, and no citation to grade), so it is not part of the shared 60-question set, the same way embeddings_search and memory are not (see docs/EVALS.md). What would actually be measured is specific to this task: given a labeled set of questions and the skill each one should trigger, selection accuracy (did it load the right one, none when none applies, and not an extra one it never used) and the token cost actually spent against what always-loading every skill would have cost. scripts/eval_run.py knows skills and refuses to score it, printing that reason; no runner for the measures above exists yet.

Run it

What to monitor

Skill selection accuracy against a labeled set of tasks, how often a loaded skill's instructions show up in the final answer, and how often the model loads a skill and then ignores it.

Cost at volume

Cost scales with how many requests actually trigger a skill load, not with how many skills exist in the registry: a large registry with rare triggers costs close to nothing per request until one is chosen.

How it fails in production

A skill checked into the project changes without review, and the next session that triggers it inherits whatever it now says or whatever tool access it now grants, with no signal that anything changed.

What to log

The full list of descriptions offered, which skill (if any) was chosen, its complete body at the time it was loaded, and which cap (if any) forced the final answer.

Try it

  1. Use it

    Find a project that ships an AGENTS.md, CLAUDE.md or a skills directory. Read one skill's body, not its description: does what it says match what the description promised?

  2. Build it

    Run python -m examples.skills --model stub:scripted from the repo root: three descriptions sit in context, the model loads warranty-checklist, and answers from its three steps. Change the name it loads in SCRIPTED (examples/skills/__main__.py) to one that does not exist: the loader reports unknown skill as text, the model answers anyway, and skills loaded reads none.

  3. Either lane

    Write two skill descriptions for a task you do regularly, worded so a person skimming them, not just a model, could tell which applies when. If you cannot tell them apart quickly, a model will not either.

  4. Either lane

    Write a SKILL.md-shaped body, a numbered procedure rather than prose, for a measurement you would otherwise explain aloud: the ripple on a switching regulator, on the oscilloscope with its 20 MHz bandwidth limit on, never on a multimeter whose AC volts function stops below the frequencies that make the ripple. Could an engineer who never watched you do it follow it? That is the test of a skill.

How it connects

Before, after and instead of this

Read first

Pages that need this one

Decoded in

Optional: products, tools, and models

5 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Apply a team’s report-writing procedure

The agent selects a report skill and loads its instructions, template, and checks when that task comes up.

Out there

Named products, tools and models

Products2
  • Antigravity CLIGoogle · coding agent · formerly Gemini CLI
  • Claude CodeAnthropic · coding agent
Tools3
  • Agent SkillsAnthropic · format for reusable agent instructions
  • AGENTS.mdopen convention · instructions file for coding agents
  • Deep AgentsLangChain · agent harness

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Agent Skills · Anthropic (accessed 09/19/2026)
  2. Extend Claude with skills · Anthropic (Claude Code documentation) (accessed 09/19/2026)
  3. AGENTS.md · agents.md (stewarded by the Agentic AI Foundation, Linux Foundation) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page