examples/synthetic_data paraphrases the site’s own 32 exact-graded questions and keeps only what
survives two checks. The first is deduplication and leakage together: a paraphrase that matches,
once case and punctuation are stripped, any of the 60 questions already in the set or any
paraphrase already accepted is rejected before it costs a second call. All 60, not just the 32 it
generates from: a paraphrase that lands on one of the 28 rubric-graded questions is a fresh copy
of a question the eval set already asks, and training on it would quietly spend the held-out value
of that question.
The second is label verification. The paraphrase is answered blind (the model sees the new
question text and nothing else, not the seed’s answer) and that answer is graded by
distillation’s grade_exact against the seed’s
accept, require and reject patterns. A paraphrase that reads fine but whose blind answer no longer
grades the same way has probably stopped asking the seed’s question, so it is dropped.
examples/synthetic_data/run.py · lines 82–163
def run(
tracer: Tracer,
model: Model,
*,
questions_path: Path = DEFAULT_QUESTIONS_PATH,
out_path: Path,
) -> SyntheticResult:
seeds: list[Question] = load_exact_questions(questions_path)
existing = all_question_texts(questions_path)
tracer.record(
kind="code",
decided_by="code",
title="Load exact-graded seed questions",
detail=f"{len(seeds)} seeds, deduplicating against {len(existing)} existing questions",
)
# Every question already in the set, so a "paraphrase" identical to the question it came from
# -- or to any other question the eval set asks, including the rubric-graded ones this example
# never generates from -- is not new data; every novel paraphrase claims its own text the
# moment it clears this check, whether or not it goes on to pass verification, so two seeds
# paraphrased the same way never both proceed. This one set is this example's dedup check and
# its leakage check at once.
seen = {_normalize(text) for text in existing}
kept: list[GeneratedQuestion] = []
rejected: list[Rejection] = []
for seed in seeds:
paraphrase = model.complete(
[Message(role="system", content=PARAPHRASE_SYSTEM_PROMPT), Message(role="user", content=seed.text)],
max_tokens=100,
)
tracer.record(
kind="model",
decided_by="code",
title="Generate a paraphrase",
detail=paraphrase.text[:200],
tokens_in=paraphrase.tokens_in,
tokens_out=paraphrase.tokens_out,
ms=paraphrase.ms,
)
normalized = _normalize(paraphrase.text)
if normalized in seen:
rejected.append(Rejection(source_id=seed.id, text=paraphrase.text, reason="duplicate"))
tracer.record(kind="code", decided_by="code", title="Reject: duplicate or unchanged", detail=seed.id)
continue
seen.add(normalized) # claim the text now: two seeds paraphrased the same way is a
# generator diversity problem whether or not this one goes on to verify
answer = model.complete(
[Message(role="system", content=ANSWER_SYSTEM_PROMPT), Message(role="user", content=paraphrase.text)],
max_tokens=200,
)
tracer.record(
kind="model",
decided_by="code",
title="Answer the paraphrase blind",
detail=answer.text[:200],
tokens_in=answer.tokens_in,
tokens_out=answer.tokens_out,
ms=answer.ms,
)
if not grade_exact(answer.text, seed):
rejected.append(Rejection(source_id=seed.id, text=paraphrase.text, reason="failed verification"))
tracer.record(
kind="code",
decided_by="code",
title="Reject: answer no longer matches the seed's grading contract",
detail=seed.id,
)
continue
kept.append(GeneratedQuestion(source_id=seed.id, text=paraphrase.text))
tracer.record(kind="code", decided_by="code", title="Keep: new and verified", detail=seed.id)
out_path.parent.mkdir(parents=True, exist_ok=True)
lines = [json.dumps({"source_id": g.source_id, "question": g.text}, sort_keys=True) for g in kept]
out_path.write_text("\n".join(lines) + ("\n" if lines else ""), encoding="utf-8", newline="\n")
tracer.record(kind="code", decided_by="code", title="Write verified synthetic questions", detail=out_path.name)
return SyntheticResult(kept=kept, rejected=rejected, out_path=out_path)
Nothing here is a model decision: the code generates, checks, asks and grades in the same order
every time, so every recorded step is decided_by: "code". Against the real set, python -m examples.synthetic_data --model stub:scripted --out .local/scratch/synthetic-data/questions.jsonl
keeps 28 and rejects 4, three for repeating text already seen and one at verification, with the
counts coming from paraphrases written down in advance rather than from a model. The same command
with --model stub keeps nothing and rejects all 32 at verification: the echoing stub’s
placeholder answer matches no real grading pattern. That is the check working on an input that was
never trying to be right.
How much it actually catches is worth being precise about, so the tests attack it. A paraphrase
that changes a number (“with the top rack removed”) is rejected, because the blind answer states
the new number and the seed’s pattern wants the old one. A negation is rejected for the same
reason: unless the answer denies the seed’s own fact in the seed’s own words, and “It does not
hold 12 place settings” contains “12 place setting”, so that one passes. That case is pinned as a
test rather than left to be discovered: this is a pattern check, not a meaning check, and a
paraphrase that quietly changed the question can still clear it. Diversity is the other blind
spot. Thirty-two paraphrases that differ from each other and reword every question the same way
grammatically pass everything here; the five question kinds in docs/EVALS.md (lookup,
multi-hop, numeric, unanswerable, conflicting sources) are the structural variety a real
generated set has to be checked for on top of text-level deduplication.
Argilla’s distilabel does this at a scale a 50-line example does not: its README calls it “the
framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable
pipelines based on verified research papers”[3], and the same README opens by saying “The
original authors have moved on to other projects” and that community collaborators have joined to
maintain it[3]: worth knowing before a pipeline depends on it.