examples/distillation runs the capture-then-filter half of the pipeline OpenAI’s guide
describes[1]: no student model is ever trained here, the same way
the fine-tuning page’s example never calls a training
API. A teacher model answers the 32 of the site’s 60 questions that are graded "exact" rather
than "rubric": a rubric question needs a grader model reading free text, which this example does
not call, so those are left out rather than approximately graded by a check they were never written
for.
grade_exact is the filter, the same accept/require/reject contract docs/EVALS.md describes for
the site’s own eval runner:
examples/distillation/run.py · lines 85–96
def grade_exact(answer: str, question: Question) -> bool:
"""The same contract `docs/EVALS.md` describes for the site's own runner: every `reject`
pattern must be absent, every `require` pattern must be present, and at least one `accept`
pattern must match when any are given. Patterns are regexes, matched case-insensitively."""
text = answer.lower()
if any(re.search(pattern, text, re.I) for pattern in question.reject):
return False
if question.require and not all(re.search(pattern, text, re.I) for pattern in question.require):
return False
if question.accept and not any(re.search(pattern, text, re.I) for pattern in question.accept):
return False
return True
run calls the teacher once per exact-graded question, grades what comes back, and writes only
what passed as chat-format JSONL: the teacher’s own words in the assistant turn, not the
question set’s answer key:
examples/distillation/run.py · lines 99–141
def run(
tracer: Tracer,
teacher: Model,
*,
questions_path: Path = DEFAULT_QUESTIONS_PATH,
out_path: Path,
) -> DistillResult:
questions = load_exact_questions(questions_path)
tracer.record(kind="code", decided_by="code", title="Load exact-graded questions", detail=f"{len(questions)} of the set")
kept: list[DistilledExample] = []
dropped: list[str] = []
for question in questions:
completion = teacher.complete(
[Message(role="system", content=TEACHER_SYSTEM_PROMPT), Message(role="user", content=question.text)],
max_tokens=200,
)
tracer.record(
kind="model",
decided_by="code",
title="Teacher answers one question",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
if grade_exact(completion.text, question):
kept.append(DistilledExample(id=question.id, question=question.text, answer=completion.text))
else:
dropped.append(question.id)
tracer.record(
kind="code",
decided_by="code",
title="Filter captured answers against the grading contract",
detail=f"{len(kept)} kept, {len(dropped)} dropped",
)
out_path.parent.mkdir(parents=True, exist_ok=True)
lines = [json.dumps(ex.as_chat_record(), sort_keys=True) for ex in kept]
out_path.write_text("\n".join(lines) + ("\n" if lines else ""), encoding="utf-8", newline="\n")
tracer.record(kind="code", decided_by="code", title="Write student training file", detail=out_path.name)
return DistillResult(kept=kept, dropped=dropped, out_path=out_path)
Every step is decided_by: "code": the code always calls the teacher, always grades the same way,
and the model’s output never changes what happens next. This is the same rule examples/rag
follows for
its own single model call. Against the real question set, python -m examples.distillation --model stub:scripted --out .local/scratch/distillation/student.jsonl keeps 30 of the 32 and names the two
it dropped, L06 and N04: a filter doing its job on a teacher that is mostly right. Those answers
are written down in advance, so the 30 is a count and not a pass rate. The same command with
--model stub keeps nothing: the echoing stub’s placeholder text matches no question’s pattern, so
all 32 are dropped. That is not a bug in the
filter; it is what an honest filter does to an answer that was never actually trying to be right,
and it is the same reason a real captured dataset needs a real teacher model before the filter’s
pass rate means anything.
Two things this example does not do, on purpose. It never checks whether a passed answer’s
reasoning was any good, only whether its final text matches a pattern: an exact-match filter is
blind to a right answer reached by a wrong method, and to the reasoning patterns DeepSeek’s paper
describes harnessing[4]. And it captures every passing answer once, with no deduplication
against near-identical phrasings; the Build it lane on
synthetic data covers the checks a larger generated
set needs and this one, at 32 questions, does not yet require.
The same shape serves an engineer whose captured data is not model answers but a log of failure
notes: 200 fault descriptions with a confirmed root cause, filtered the way grade_exact filters
an answer here, then used to train a small classifier that tags a new note with a likely category.
Whether 200 is enough is not a number this page can give; it is the same before/after question the
eval section below asks of any claim of improvement, in any of the three settings the notes came
from. A triage classifier reading a production line’s daily failure log is scored against a slice
of that log’s own history withheld from training. A classifier trained on a handful of bring-up
notes from engineering test is scored the same way, on fewer notes, with a correspondingly smaller
claim. A classifier meant to flag a note worth a second look before a measurement ships is scored
hardest of all, since what it feeds is a person’s decision to trust a number, and its own output is
never the verdict.