# Distillation

_Topics at every level · sourced_

Training a smaller model to reproduce what a larger one does on your task.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a larger system's outputs into a candidate smaller model and an independent check. Inspect which useful behavior survives and which teacher errors can be copied.

**Assumptions:** Teacher outputs are proposed training material, not ground truth. The student may operate with different capacity and context constraints.

**Design choices:** Select examples that cover the intended workload, review consequential labels, and measure the student directly against requirements and a baseline.

**Request:** Design a smaller classifier from a larger model's reviewed labels.

**Starting evidence:** Teacher labels 100 fictional examples; audit finds five errors. Independent evaluation set exists.

**Action and control:** Correct or exclude teacher mistakes before training; evaluate the student independently.

**Stage records (authored, not executed):**

### Input record

Teacher labels 100 fictional examples; audit finds five errors. Independent evaluation set exists.

What changed: Establish the facts supplied for this version of the task.

### Design note

Select examples that cover the intended workload, review consequential labels, and measure the student directly against requirements and a baseline.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Correct or exclude teacher mistakes before training; evaluate the student independently.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Reviewed label set prepared. Quality and resource tradeoffs await real measurement.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Teacher labels, human corrections, separate evaluation set, and a labeled illustrative quality/resource tradeoff.

If the result falls short:
If the student copies a systematic error or loses rare-case performance, correct the dataset or narrow its role. Matching average teacher behavior may be insufficient.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for a bounded classifier or other repeated task. Decide which quality, latency, and resource tradeoffs are acceptable before judging the smaller model.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Reviewed label set prepared. Quality and resource tradeoffs await real measurement.

**Change something — Trust every teacher label:** Mistakes become training targets. Teacher agreement is not task correctness.

**Decision:** Is matching the teacher sufficient evidence?

**Answer:** No; use independent ground truth.

**Why:** Teacher errors transfer to the student; lower cost can come with reduced coverage or calibration.

**Review criteria:** Teacher labels, human corrections, separate evaluation set, and a labeled illustrative quality/resource tradeoff.

**Recovery:** If the student copies a systematic error or loses rare-case performance, correct the dataset or narrow its role. Matching average teacher behavior may be insufficient.

**Adapt it:** Use this for a bounded classifier or other repeated task. Decide which quality, latency, and resource tradeoffs are acceptable before judging the smaller model.

A larger teacher model produces material; a smaller student model is trained on that instead of on
data a person wrote. That is distillation, and it belongs to
[changing the model](/gradient_ascent/techniques/adaptation/) alongside
[fine-tuning](/gradient_ascent/techniques/fine-tuning/), which is what the student's training run
actually is once the dataset exists.

OpenAI's own distillation guide lays out a four-step flow; the middle two are the mechanism this
page and its example build: "Capture results generated from your model" and then "Use the captured
responses from the large model that fit your criteria to generate a dataset"[1]: the
teacher's outputs, filtered before anything is trained on them. What gets captured need not stop
at the final answer. DeepSeek's paper on its R1 model
reports that "the emergent reasoning patterns exhibited by these large-scale models can be
systematically harnessed to guide and enhance the reasoning capabilities of smaller
models"[4]: what that paper describes carrying over is reasoning patterns, not a list of
conclusions.

This page is sourced, not measured: distillation is described from primary sources, but no
training run has happened here, and the example below stops exactly where a real project would
start paying for one.

## Practical guidance

A small, fast model marketed as unusually good at one narrow job may be distilled: a maker ran a
larger model over many examples of that job and trained a smaller model on the results. Before
signing anything built that way, send whoever is selling it one procurement question in writing:
"Was this model trained on outputs captured from another company's model, and does that company's
terms allow training a model you resell to us on those outputs?" This site gives no legal advice
and cannot tell you how a clause applies to your plan; it can quote three documents as they read
today, so you know what to ask a vendor to explain.

Anthropic's Commercial Terms of Service state under Use Restrictions that "Customer may not and
must not attempt to (a) access the Services to build a competing product or service, including to
train competing AI models or resell the Services except as expressly approved by
Anthropic"[2]. Google's Gemini API Additional Terms of Service state, "You may not use
the Services to develop models that compete with the Services (e.g., Gemini API or Google AI
Studio)"[3], while separately saying "Google only uses content that you import or upload
to our model tuning feature for that express purpose"[3], a statement about data use, not
an exception to the restriction above.

OpenAI's Services Agreement restricts a customer, "except for a Permitted Exception," from using
"Output to develop artificial intelligence models that compete with OpenAI's products and
services"[5]. That exception covers Output used to "develop artificial intelligence
models primarily intended to categorize, classify, or organize data (e.g., embeddings or
classifiers), if these models are not distributed or made commercially available to third
parties," and to fine tune or customize "models provided as part of OpenAI's fine-tuning or other
Services"[5]: an in-house classifier fits; a model sold to a third party does not.

Whether any of that covers your actual plan is a question for whoever can read your contract, not
this page. Ask a second, technical question alongside the legal one: what task were the captured
answers filtered for, and does your use fall inside it or outside it? A model distilled on support
replies for one product answers a question about a different one fluently and wrong, with nothing
in the reply flagging that it has left the task it was trained for.

## Implementation details

`examples/distillation` runs the capture-then-filter half of the pipeline OpenAI's guide
describes[1]: no student model is ever trained here, the same way
[the fine-tuning page's](/gradient_ascent/techniques/fine-tuning/) example never calls a training
API. A teacher model answers the 32 of the site's 60 questions that are graded `"exact"` rather
than `"rubric"`: a rubric question needs a grader model reading free text, which this example does
not call, so those are left out rather than approximately graded by a check they were never written
for.

`grade_exact` is the filter, the same accept/require/reject contract `docs/EVALS.md` describes for
the site's own eval runner:

`examples/distillation/run.py` (lines 85-96)

```python
def grade_exact(answer: str, question: Question) -> bool:
    """The same contract `docs/EVALS.md` describes for the site's own runner: every `reject`
    pattern must be absent, every `require` pattern must be present, and at least one `accept`
    pattern must match when any are given. Patterns are regexes, matched case-insensitively."""
    text = answer.lower()
    if any(re.search(pattern, text, re.I) for pattern in question.reject):
        return False
    if question.require and not all(re.search(pattern, text, re.I) for pattern in question.require):
        return False
    if question.accept and not any(re.search(pattern, text, re.I) for pattern in question.accept):
        return False
    return True
```

`run` calls the teacher once per exact-graded question, grades what comes back, and writes only
what passed as chat-format JSONL: the teacher's own words in the assistant turn, not the
question set's answer key:

`examples/distillation/run.py` (lines 99-141)

```python
def run(
    tracer: Tracer,
    teacher: Model,
    *,
    questions_path: Path = DEFAULT_QUESTIONS_PATH,
    out_path: Path,
) -> DistillResult:
    questions = load_exact_questions(questions_path)
    tracer.record(kind="code", decided_by="code", title="Load exact-graded questions", detail=f"{len(questions)} of the set")

    kept: list[DistilledExample] = []
    dropped: list[str] = []
    for question in questions:
        completion = teacher.complete(
            [Message(role="system", content=TEACHER_SYSTEM_PROMPT), Message(role="user", content=question.text)],
            max_tokens=200,
        )
        tracer.record(
            kind="model",
            decided_by="code",
            title="Teacher answers one question",
            detail=completion.text[:200],
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
        if grade_exact(completion.text, question):
            kept.append(DistilledExample(id=question.id, question=question.text, answer=completion.text))
        else:
            dropped.append(question.id)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Filter captured answers against the grading contract",
        detail=f"{len(kept)} kept, {len(dropped)} dropped",
    )

    out_path.parent.mkdir(parents=True, exist_ok=True)
    lines = [json.dumps(ex.as_chat_record(), sort_keys=True) for ex in kept]
    out_path.write_text("\n".join(lines) + ("\n" if lines else ""), encoding="utf-8", newline="\n")
    tracer.record(kind="code", decided_by="code", title="Write student training file", detail=out_path.name)

    return DistillResult(kept=kept, dropped=dropped, out_path=out_path)
```

Every step is `decided_by: "code"`: the code always calls the teacher, always grades the same way,
and the model's output never changes what happens next. This is the same rule `examples/rag`
follows for
its own single model call. Against the real question set, `python -m examples.distillation --model
stub:scripted --out .local/scratch/distillation/student.jsonl` keeps 30 of the 32 and names the two
it dropped, L06 and N04: a filter doing its job on a teacher that is mostly right. Those answers
are written down in advance, so the 30 is a count and not a pass rate. The same command with
`--model stub` keeps nothing: the echoing stub's placeholder text matches no question's pattern, so
all 32 are dropped. That is not a bug in the
filter; it is what an honest filter does to an answer that was never actually trying to be right,
and it is the same reason a real captured dataset needs a real teacher model before the filter's
pass rate means anything.

Two things this example does not do, on purpose. It never checks whether a passed answer's
reasoning was any good, only whether its final text matches a pattern: an exact-match filter is
blind to a right answer reached by a wrong method, and to the reasoning patterns DeepSeek's paper
describes harnessing[4]. And it captures every passing answer once, with no deduplication
against near-identical phrasings; the Build it lane on
[synthetic data](/gradient_ascent/techniques/synthetic-data/) covers the checks a larger generated
set needs and this one, at 32 questions, does not yet require.

The same shape serves an engineer whose captured data is not model answers but a log of failure
notes: 200 fault descriptions with a confirmed root cause, filtered the way `grade_exact` filters
an answer here, then used to train a small classifier that tags a new note with a likely category.
Whether 200 is enough is not a number this page can give; it is the same before/after question the
eval section below asks of any claim of improvement, in any of the three settings the notes came
from. A triage classifier reading a production line's daily failure log is scored against a slice
of that log's own history withheld from training. A classifier trained on a handful of bring-up
notes from engineering test is scored the same way, on fewer notes, with a correspondingly smaller
claim. A classifier meant to flag a note worth a second look before a measurement ships is scored
hardest of all, since what it feeds is a person's decision to trust a number, and its own output is
never the verdict.

## When you do not need this

Try the teacher model itself, with a good prompt, before distilling anything from it. If a
well-written prompt against the larger model already gets the accuracy and consistency you need,
training a smaller model on its outputs adds a dataset to build, a filter to trust, and a training
run to pay for, in exchange for a cost saving you have not yet shown you need.

Distillation earns its cost once the larger model's per-call price or latency, multiplied by real
call volume, is the actual problem, not before. A task called a few times a day rarely justifies
building and maintaining a captured, filtered dataset just to run it on cheaper hardware.

## Failure modes

### A shallow filter passes a right-looking wrong answer

- **How to notice it:** A captured answer matches the exact-match pattern the way the filter shown on this page checks it, but is wrong for a reason the pattern was never built to catch: the right number attached to the wrong appliance, say.
- **How to test for it:** Hand-read a sample of what the filter kept, not just its pass rate. A pattern check only ever tests what its author thought to write a pattern for.

### The student inherits the teacher’s confident mistakes

- **How to notice it:** The teacher model is systematically wrong about one thing, every captured answer about it reads fluently and passes the filter, and the student learns the same wrong answer, now delivered faster and cheaper.
- **How to test for it:** Before training on a captured set, check the teacher's own accuracy on a sample graded by a person, not only by the pattern filter this page's example uses.

### Narrow capture mistaken for broad capability

- **How to notice it:** A student distilled on one task's captured answers performs well on that task and confidently wrong outside it, in the same way a fine-tuned model does, because nothing about distillation preserves what the teacher could do beyond what was captured.
- **How to test for it:** Ask the student a question clearly outside the captured task and compare its answer against the teacher's own; a gap that only shows up outside the task is this failure.

### Captured outputs used without reading the terms

- **How to notice it:** A team builds and ships a product trained on a hosted model's captured outputs, and nobody has read what that maker's current terms say about training models on them. All three makers quoted on this page carry a clause about competing models, each with its own scope and its own exceptions.
- **How to test for it:** Before capturing anything, open the current terms of the maker you are actually using and find the use-restriction section. Whether your plan falls inside a clause is a question for someone who can advise on it, not for a technique page.

### No filter at all for a rubric-graded task

- **How to notice it:** A captured dataset for an open-ended task has no exact-match pattern to filter by, so everything the teacher produced goes into training unfiltered, including answers a person would have rejected.
- **How to test for it:** Check whether every kept example passed some check, even a cheap one, before training on it; 'the teacher produced it' is not a filter.

## How to Evaluate It

The example answers questions from the teacher model's own knowledge, with no retrieval and no
citations, so the site's 60-question grading contract (which checks citations against the corpus)
has nothing to grade it on (see `docs/EVALS.md`). The number it does produce is its filter's pass
rate over the 32 exact-graded questions, and that number is not accuracy: a question with one short
accept pattern is easier to pass than one carrying several `require` patterns, so the rate reflects
how the patterns were written as much as how good the teacher was.

The measurement that would settle anything happens after training, not during capture. Score the
finished student on the same 60 questions the site runs against every other technique, against the
same questions run on whatever it replaced, and report both. A pass rate collected while building
the dataset is not a result about the student.

## Run it

**What to monitor.** The captured dataset's pass rate against the filter over time, and, on a sample, whether
  the teacher's own answers were actually right: a filter checks the pattern, not the fact.

**Cost at volume.** Capturing and filtering is a one-time or periodic cost that scales with how many
  examples you capture, not with how many times the student answers afterward; the student's
  own per-call cost is what should fall once it is trained and serving real traffic.

**How it fails in production.** The teacher model the dataset was captured from is replaced or updated by its
  maker, and the student, trained on the old teacher's answers, keeps giving the old teacher's
  answer to a question the new teacher would now answer differently.

**What to log.** The teacher model id and the date it was captured, the filter's pass rate, and the
  training data's version, so a question about the student's behavior can be traced to which
  teacher, and which filtered set, produced it.

## Try it

1. **Use it.** Find a small model marketed as distilled from a larger one. Check the maker's own page for what task the distillation covered, then try it on something outside that task.
2. **Build it.** Run python -m examples.distillation --model stub:scripted --out .local/scratch/distillation/student.jsonl from the repo root and read the dropped list it prints, L06 and N04, a made-up error code and an arithmetic slip. Then open examples/distillation/run.py and change TEACHER_SYSTEM_PROMPT to ask for a one-word answer instead of a sentence; against a real teacher, would the pass rate go up or down, and why?
3. **Either lane.** Take one exact-graded question from evals/questions.json and write two answers by hand: one factually right that fails grade_exact, one factually wrong that passes. Both being constructible is the filter's blind spot, not a bug.


## Sources

1. [Supervised fine-tuning](https://developers.openai.com/api/docs/guides/supervised-fine-tuning#distilling-from-a-larger-model) — OpenAI (API documentation) (accessed 2026-09-19)
2. [Commercial Terms of Service](https://www.anthropic.com/legal/commercial-terms) — Anthropic, 2025-06-17 (accessed 2026-09-19)
3. [Gemini API Additional Terms of Service](https://ai.google.dev/gemini-api/terms) — Google, 2026-03-23 (accessed 2026-09-19)
4. [DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning](https://arxiv.org/abs/2501.12948) — arXiv (DeepSeek-AI), 2025-01-22 (accessed 2026-09-19)
5. [OpenAI Services Agreement](https://openai.com/policies/business-terms/) — OpenAI, 2026-01-01 (accessed 2026-09-19)


Last reviewed 2026-09-19.
