Topics at every level

Distillation

Training a smaller model to reproduce what a larger one does on your task.

Sourced

Concept at a glance

Teach a smaller model from a larger model’s work.

SequenceConceptual illustration
Teach a smaller model from a larger model’s work.Teacher model leads to Check the examples. Check the examples leads to Student model. The student must be evaluated on the task, not just on agreement with the teacher.Teacher modelGenerate demonstrationsCheck the examplesFilter the training materialStudent modelTrain and evaluateTeach a smaller model from a larger model’s work.Teacher model leads to Check the examples. Check the examples leads to Student model. The student must be evaluated on the task, not just on agreement with the teacher.Teacher modelGenerate demonstrationsCheck the examplesFilter the training materialStudent modelTrain and evaluate
Read the connections in words
  • Teacher model → Check the examples: Filter the training material.
  • Check the examples → Student model: Train and evaluate.
Key idea

The student must be evaluated on the task, not just on agreement with the teacher.

A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Distillation: see it in practice.

Training a student model to approximate selected behavior of a teacher model.

What you’ll walk through

Follow a larger system's outputs into a candidate smaller model and an independent check. Inspect which useful behavior survives and which teacher errors can be copied.

The task in this version

Design a smaller classifier from a larger model's reviewed labels.

What you’ll learn to check

Teacher labels, human corrections, separate evaluation set, and a labeled illustrative quality/resource tradeoff.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Design a smaller classifier from a larger model's reviewed labels.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Teacher labels 100 fictional examples; audit finds five errors. Independent evaluation set exists.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Teacher outputs are proposed training material, not ground truth. The student may operate with different capacity and context constraints.

1 / 6

Apply this to your project

Describe your task to your own model and use Distillation as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

A larger teacher model produces material; a smaller student model is trained on that instead of on data a person wrote. That is distillation, and it belongs to changing the model alongside fine-tuning, which is what the student’s training run actually is once the dataset exists.

OpenAI’s own distillation guide lays out a four-step flow; the middle two are the mechanism this page and its example build: “Capture results generated from your model” and then “Use the captured responses from the large model that fit your criteria to generate a dataset”[1]: the teacher’s outputs, filtered before anything is trained on them. What gets captured need not stop at the final answer. DeepSeek’s paper on its R1 model reports that “the emergent reasoning patterns exhibited by these large-scale models can be systematically harnessed to guide and enhance the reasoning capabilities of smaller models”[4]: what that paper describes carrying over is reasoning patterns, not a list of conclusions.

This page is sourced, not measured: distillation is described from primary sources, but no training run has happened here, and the example below stops exactly where a real project would start paying for one.

Practical guidance

A small, fast model marketed as unusually good at one narrow job may be distilled: a maker ran a larger model over many examples of that job and trained a smaller model on the results. Before signing anything built that way, send whoever is selling it one procurement question in writing: “Was this model trained on outputs captured from another company’s model, and does that company’s terms allow training a model you resell to us on those outputs?” This site gives no legal advice and cannot tell you how a clause applies to your plan; it can quote three documents as they read today, so you know what to ask a vendor to explain.

Anthropic’s Commercial Terms of Service state under Use Restrictions that “Customer may not and must not attempt to (a) access the Services to build a competing product or service, including to train competing AI models or resell the Services except as expressly approved by Anthropic”[2]. Google’s Gemini API Additional Terms of Service state, “You may not use the Services to develop models that compete with the Services (e.g., Gemini API or Google AI Studio)”[3], while separately saying “Google only uses content that you import or upload to our model tuning feature for that express purpose”[3], a statement about data use, not an exception to the restriction above.

OpenAI’s Services Agreement restricts a customer, “except for a Permitted Exception,” from using “Output to develop artificial intelligence models that compete with OpenAI’s products and services”[5]. That exception covers Output used to “develop artificial intelligence models primarily intended to categorize, classify, or organize data (e.g., embeddings or classifiers), if these models are not distributed or made commercially available to third parties,” and to fine tune or customize “models provided as part of OpenAI’s fine-tuning or other Services”[5]: an in-house classifier fits; a model sold to a third party does not.

Whether any of that covers your actual plan is a question for whoever can read your contract, not this page. Ask a second, technical question alongside the legal one: what task were the captured answers filtered for, and does your use fall inside it or outside it? A model distilled on support replies for one product answers a question about a different one fluently and wrong, with nothing in the reply flagging that it has left the task it was trained for.

Implementation details

examples/distillation runs the capture-then-filter half of the pipeline OpenAI’s guide describes[1]: no student model is ever trained here, the same way the fine-tuning page’s example never calls a training API. A teacher model answers the 32 of the site’s 60 questions that are graded "exact" rather than "rubric": a rubric question needs a grader model reading free text, which this example does not call, so those are left out rather than approximately graded by a check they were never written for.

grade_exact is the filter, the same accept/require/reject contract docs/EVALS.md describes for the site’s own eval runner:

examples/distillation/run.py · lines 85–96
def grade_exact(answer: str, question: Question) -> bool:
    """The same contract `docs/EVALS.md` describes for the site's own runner: every `reject`
    pattern must be absent, every `require` pattern must be present, and at least one `accept`
    pattern must match when any are given. Patterns are regexes, matched case-insensitively."""
    text = answer.lower()
    if any(re.search(pattern, text, re.I) for pattern in question.reject):
        return False
    if question.require and not all(re.search(pattern, text, re.I) for pattern in question.require):
        return False
    if question.accept and not any(re.search(pattern, text, re.I) for pattern in question.accept):
        return False
    return True

run calls the teacher once per exact-graded question, grades what comes back, and writes only what passed as chat-format JSONL: the teacher’s own words in the assistant turn, not the question set’s answer key:

examples/distillation/run.py · lines 99–141
def run(
    tracer: Tracer,
    teacher: Model,
    *,
    questions_path: Path = DEFAULT_QUESTIONS_PATH,
    out_path: Path,
) -> DistillResult:
    questions = load_exact_questions(questions_path)
    tracer.record(kind="code", decided_by="code", title="Load exact-graded questions", detail=f"{len(questions)} of the set")

    kept: list[DistilledExample] = []
    dropped: list[str] = []
    for question in questions:
        completion = teacher.complete(
            [Message(role="system", content=TEACHER_SYSTEM_PROMPT), Message(role="user", content=question.text)],
            max_tokens=200,
        )
        tracer.record(
            kind="model",
            decided_by="code",
            title="Teacher answers one question",
            detail=completion.text[:200],
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
        if grade_exact(completion.text, question):
            kept.append(DistilledExample(id=question.id, question=question.text, answer=completion.text))
        else:
            dropped.append(question.id)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Filter captured answers against the grading contract",
        detail=f"{len(kept)} kept, {len(dropped)} dropped",
    )

    out_path.parent.mkdir(parents=True, exist_ok=True)
    lines = [json.dumps(ex.as_chat_record(), sort_keys=True) for ex in kept]
    out_path.write_text("\n".join(lines) + ("\n" if lines else ""), encoding="utf-8", newline="\n")
    tracer.record(kind="code", decided_by="code", title="Write student training file", detail=out_path.name)

    return DistillResult(kept=kept, dropped=dropped, out_path=out_path)

Every step is decided_by: "code": the code always calls the teacher, always grades the same way, and the model’s output never changes what happens next. This is the same rule examples/rag follows for its own single model call. Against the real question set, python -m examples.distillation --model stub:scripted --out .local/scratch/distillation/student.jsonl keeps 30 of the 32 and names the two it dropped, L06 and N04: a filter doing its job on a teacher that is mostly right. Those answers are written down in advance, so the 30 is a count and not a pass rate. The same command with --model stub keeps nothing: the echoing stub’s placeholder text matches no question’s pattern, so all 32 are dropped. That is not a bug in the filter; it is what an honest filter does to an answer that was never actually trying to be right, and it is the same reason a real captured dataset needs a real teacher model before the filter’s pass rate means anything.

Two things this example does not do, on purpose. It never checks whether a passed answer’s reasoning was any good, only whether its final text matches a pattern: an exact-match filter is blind to a right answer reached by a wrong method, and to the reasoning patterns DeepSeek’s paper describes harnessing[4]. And it captures every passing answer once, with no deduplication against near-identical phrasings; the Build it lane on synthetic data covers the checks a larger generated set needs and this one, at 32 questions, does not yet require.

The same shape serves an engineer whose captured data is not model answers but a log of failure notes: 200 fault descriptions with a confirmed root cause, filtered the way grade_exact filters an answer here, then used to train a small classifier that tags a new note with a likely category. Whether 200 is enough is not a number this page can give; it is the same before/after question the eval section below asks of any claim of improvement, in any of the three settings the notes came from. A triage classifier reading a production line’s daily failure log is scored against a slice of that log’s own history withheld from training. A classifier trained on a handful of bring-up notes from engineering test is scored the same way, on fewer notes, with a correspondingly smaller claim. A classifier meant to flag a note worth a second look before a measurement ships is scored hardest of all, since what it feeds is a person’s decision to trust a number, and its own output is never the verdict.

When you do not need this

Try the teacher model itself, with a good prompt, before distilling anything from it. If a well-written prompt against the larger model already gets the accuracy and consistency you need, training a smaller model on its outputs adds a dataset to build, a filter to trust, and a training run to pay for, in exchange for a cost saving you have not yet shown you need.

Distillation earns its cost once the larger model’s per-call price or latency, multiplied by real call volume, is the actual problem, not before. A task called a few times a day rarely justifies building and maintaining a captured, filtered dataset just to run it on cheaper hardware.

Failure modes

A shallow filter passes a right-looking wrong answer

How to notice it
A captured answer matches the exact-match pattern the way the filter shown on this page checks it, but is wrong for a reason the pattern was never built to catch: the right number attached to the wrong appliance, say.
How to test for it
Hand-read a sample of what the filter kept, not just its pass rate. A pattern check only ever tests what its author thought to write a pattern for.

The student inherits the teacher’s confident mistakes

How to notice it
The teacher model is systematically wrong about one thing, every captured answer about it reads fluently and passes the filter, and the student learns the same wrong answer, now delivered faster and cheaper.
How to test for it
Before training on a captured set, check the teacher's own accuracy on a sample graded by a person, not only by the pattern filter this page's example uses.

Narrow capture mistaken for broad capability

How to notice it
A student distilled on one task's captured answers performs well on that task and confidently wrong outside it, in the same way a fine-tuned model does, because nothing about distillation preserves what the teacher could do beyond what was captured.
How to test for it
Ask the student a question clearly outside the captured task and compare its answer against the teacher's own; a gap that only shows up outside the task is this failure.

Captured outputs used without reading the terms

How to notice it
A team builds and ships a product trained on a hosted model's captured outputs, and nobody has read what that maker's current terms say about training models on them. All three makers quoted on this page carry a clause about competing models, each with its own scope and its own exceptions.
How to test for it
Before capturing anything, open the current terms of the maker you are actually using and find the use-restriction section. Whether your plan falls inside a clause is a question for someone who can advise on it, not for a technique page.

No filter at all for a rubric-graded task

How to notice it
A captured dataset for an open-ended task has no exact-match pattern to filter by, so everything the teacher produced goes into training unfiltered, including answers a person would have rejected.
How to test for it
Check whether every kept example passed some check, even a cheap one, before training on it; 'the teacher produced it' is not a filter.

How to Evaluate It

The example answers questions from the teacher model’s own knowledge, with no retrieval and no citations, so the site’s 60-question grading contract (which checks citations against the corpus) has nothing to grade it on (see docs/EVALS.md). The number it does produce is its filter’s pass rate over the 32 exact-graded questions, and that number is not accuracy: a question with one short accept pattern is easier to pass than one carrying several require patterns, so the rate reflects how the patterns were written as much as how good the teacher was.

The measurement that would settle anything happens after training, not during capture. Score the finished student on the same 60 questions the site runs against every other technique, against the same questions run on whatever it replaced, and report both. A pass rate collected while building the dataset is not a result about the student.

Run it

What to monitor

The captured dataset's pass rate against the filter over time, and, on a sample, whether the teacher's own answers were actually right: a filter checks the pattern, not the fact.

Cost at volume

Capturing and filtering is a one-time or periodic cost that scales with how many examples you capture, not with how many times the student answers afterward; the student's own per-call cost is what should fall once it is trained and serving real traffic.

How it fails in production

The teacher model the dataset was captured from is replaced or updated by its maker, and the student, trained on the old teacher's answers, keeps giving the old teacher's answer to a question the new teacher would now answer differently.

What to log

The teacher model id and the date it was captured, the filter's pass rate, and the training data's version, so a question about the student's behavior can be traced to which teacher, and which filtered set, produced it.

Try it

  1. Use it

    Find a small model marketed as distilled from a larger one. Check the maker's own page for what task the distillation covered, then try it on something outside that task.

  2. Build it

    Run python -m examples.distillation --model stub:scripted --out .local/scratch/distillation/student.jsonl from the repo root and read the dropped list it prints, L06 and N04, a made-up error code and an arithmetic slip. Then open examples/distillation/run.py and change TEACHER_SYSTEM_PROMPT to ask for a one-word answer instead of a sentence; against a real teacher, would the pass rate go up or down, and why?

  3. Either lane

    Take one exact-graded question from evals/questions.json and write two answers by hand: one factually right that fails grade_exact, one factually wrong that passes. Both being constructible is the filter's blind spot, not a bug.

How it connects

Before, after and instead of this

Optional: products, tools, and models

Concrete examples

In practice

Train a smaller task specialist

Have a larger teacher generate demonstrations, verify them, and train a smaller student on the accepted examples.

An illustrative task example. No verified product or tool is currently listed for this concept.

Out there

Named products, tools and models

No product, tool or model is registered against this page yet. The names index lists every one the site does name, and which technique each belongs to.

Open the names index →

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Supervised fine-tuning · OpenAI (API documentation) (accessed 09/19/2026)
  2. Commercial Terms of Service · Anthropic, 06/17/2025 (accessed 09/19/2026)
  3. Gemini API Additional Terms of Service · Google, 03/23/2026 (accessed 09/19/2026)
  4. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI), 01/22/2025 (accessed 09/19/2026)
  5. OpenAI Services Agreement · OpenAI, 01/01/2026 (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page