# Synthetic data

_Topics at every level · sourced_

Using a model to write training or test examples, and checking them before they are used.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a gap in an example collection through generated candidates, filtering, and a check on independent data. Inspect whether the new examples add useful variation or repeat the generator's assumptions.

**Assumptions:** Generated cases can be repetitive, unrealistic, or mislabeled. They may miss precisely the unusual situations the real task contains.

**Design choices:** Use synthetic data to complement evidence where justified, with review and clear provenance. Keep real or independently constructed evaluation cases separate.

**Request:** Expand rare support categories with realistic examples.

**Starting evidence:** Batch repeats address-change wording with different names and includes mislabeled cancellations.

**Action and control:** Review labels, deduplicate patterns, and compare coverage with real messages.

**Stage records (authored, not executed):**

### Input record

Batch repeats address-change wording with different names and includes mislabeled cancellations.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use synthetic data to complement evidence where justified, with review and clear provenance. Keep real or independently constructed evaluation cases separate.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Review labels, deduplicate patterns, and compare coverage with real messages.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Keep distinct correct examples; reject cancellations and near-duplicates. Evaluate on independent real cases.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Generation brief, accepted/rejected examples, diversity checks, and evaluation on independently collected real cases.

If the result falls short:
If apparent gains disappear on independent cases, inspect duplicates, leakage, and unrealistic patterns. Generate from a revised coverage plan rather than merely increasing volume.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to practice cases, extraction variants, or rare categories. Define the missing coverage and how candidate examples will be accepted or rejected.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Keep distinct correct examples; reject cancellations and near-duplicates. Evaluate on independent real cases.

**Change something — Use the same generated batch for training and test:** Success is circular and does not establish real-message performance.

**Decision:** Does a larger generated dataset always improve coverage?

**Answer:** No; diversity and correctness need review.

**Why:** Duplicates, unrealistic language, and label errors can make a dataset look larger without adding useful coverage.

**Review criteria:** Generation brief, accepted/rejected examples, diversity checks, and evaluation on independently collected real cases.

**Recovery:** If apparent gains disappear on independent cases, inspect duplicates, leakage, and unrealistic patterns. Generate from a revised coverage plan rather than merely increasing volume.

**Adapt it:** Apply this to practice cases, extraction variants, or rare categories. Define the missing coverage and how candidate examples will be accepted or rejected.

A model writes the training or test examples instead of a person. That is synthetic data; it sits
under [changing the model](/gradient_ascent/techniques/adaptation/), and it is where the examples
[fine-tuning](/gradient_ascent/techniques/fine-tuning/) and
[distillation](/gradient_ascent/techniques/distillation/) need can come from when nobody has time
to write them.

Generation is half of it. The Self-Instruct paper puts its own pipeline in one sentence: "Our
pipeline generates instructions, input, and output samples from a language model, then filters
invalid or similar ones before using them to finetune the original model"[1]: generate,
then filter, as two steps rather than one. The filter is the half that gets skipped. A 2023 paper
titled *The Curse of Recursion* asks what becomes of a model "once LLMs contribute much of the
language found online": "We find that use of model-generated content in training causes
irreversible defects in the resulting models, where tails of the original content distribution
disappear. We refer to this effect as Model Collapse and show that it can occur in Variational
Autoencoders, Gaussian Mixture Models and LLMs."[2] A finding about web-scale generated
content, shown in three named settings, not a verdict on every dataset a model helped write.

This page is sourced, not measured: the generation and filtering methods below come from their
authors' own papers, and no generated dataset has been used for anything on this site.

## Practical guidance

Ask one question of any product or paper that reports a dataset built partly or fully by a model,
and put it exactly this way: "What checked each example before it was used, and what share did
that check reject?" A number stated as "50,000 synthetic examples" says nothing on its own about
whether a person, or any independent process, verified a single one of them. Generating and
filtering are two separate steps, the way Self-Instruct's own pipeline builds a filtering step in
on purpose rather than treating generation as the whole job[1]; a dataset description
that only ever mentions the first step, with no rejection rate anywhere, is answering a question
you did not ask.

"Model collapse" is not a reason to distrust every generated example. It names one specific
failure: a model's own unchecked output feeding the next model's training, round after round, so
that the rare and unusual cases quietly vanish from what later models ever see[2]. One
generated batch, checked once against something outside the model that produced it, and used once,
is not that loop. Ask what the check actually compared the generated data against, not just
whether the word "filtered" appears on the page.

None of this is specific to text. The same question applies to a generated image set, a generated
audio set, or a set of generated tool-call examples: checked against something outside the model
that made it, or only against how plausible it looks.

If the claim in front of you is about a measurement rather than words, the question does not need
asking: refuse it outright. A model may help draft the report around a reading, a margin, or a
pass or fail line; it may not produce the reading itself. Five boards back from a supplier is not
enough characterization data, and no amount of generated language changes that. Generating more
fault descriptions or operator notes to train or test a classifier is a different, honest use of
the same idea, so long as the notes are real language checked before use, not a stand-in for the
readings themselves.

## Implementation details

`examples/synthetic_data` paraphrases the site's own 32 exact-graded questions and keeps only what
survives two checks. The first is deduplication and leakage together: a paraphrase that matches,
once case and punctuation are stripped, any of the 60 questions already in the set or any
paraphrase already accepted is rejected before it costs a second call. All 60, not just the 32 it
generates from: a paraphrase that lands on one of the 28 rubric-graded questions is a fresh copy
of a question the eval set already asks, and training on it would quietly spend the held-out value
of that question.

The second is label verification. The paraphrase is answered blind (the model sees the new
question text and nothing else, not the seed's answer) and that answer is graded by
[distillation](/gradient_ascent/techniques/distillation/)'s `grade_exact` against the *seed's*
accept, require and reject patterns. A paraphrase that reads fine but whose blind answer no longer
grades the same way has probably stopped asking the seed's question, so it is dropped.

`examples/synthetic_data/run.py` (lines 82-163)

```python
def run(
    tracer: Tracer,
    model: Model,
    *,
    questions_path: Path = DEFAULT_QUESTIONS_PATH,
    out_path: Path,
) -> SyntheticResult:
    seeds: list[Question] = load_exact_questions(questions_path)
    existing = all_question_texts(questions_path)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Load exact-graded seed questions",
        detail=f"{len(seeds)} seeds, deduplicating against {len(existing)} existing questions",
    )

    # Every question already in the set, so a "paraphrase" identical to the question it came from
    # -- or to any other question the eval set asks, including the rubric-graded ones this example
    # never generates from -- is not new data; every novel paraphrase claims its own text the
    # moment it clears this check, whether or not it goes on to pass verification, so two seeds
    # paraphrased the same way never both proceed. This one set is this example's dedup check and
    # its leakage check at once.
    seen = {_normalize(text) for text in existing}
    kept: list[GeneratedQuestion] = []
    rejected: list[Rejection] = []

    for seed in seeds:
        paraphrase = model.complete(
            [Message(role="system", content=PARAPHRASE_SYSTEM_PROMPT), Message(role="user", content=seed.text)],
            max_tokens=100,
        )
        tracer.record(
            kind="model",
            decided_by="code",
            title="Generate a paraphrase",
            detail=paraphrase.text[:200],
            tokens_in=paraphrase.tokens_in,
            tokens_out=paraphrase.tokens_out,
            ms=paraphrase.ms,
        )

        normalized = _normalize(paraphrase.text)
        if normalized in seen:
            rejected.append(Rejection(source_id=seed.id, text=paraphrase.text, reason="duplicate"))
            tracer.record(kind="code", decided_by="code", title="Reject: duplicate or unchanged", detail=seed.id)
            continue
        seen.add(normalized)  # claim the text now: two seeds paraphrased the same way is a
        # generator diversity problem whether or not this one goes on to verify

        answer = model.complete(
            [Message(role="system", content=ANSWER_SYSTEM_PROMPT), Message(role="user", content=paraphrase.text)],
            max_tokens=200,
        )
        tracer.record(
            kind="model",
            decided_by="code",
            title="Answer the paraphrase blind",
            detail=answer.text[:200],
            tokens_in=answer.tokens_in,
            tokens_out=answer.tokens_out,
            ms=answer.ms,
        )

        if not grade_exact(answer.text, seed):
            rejected.append(Rejection(source_id=seed.id, text=paraphrase.text, reason="failed verification"))
            tracer.record(
                kind="code",
                decided_by="code",
                title="Reject: answer no longer matches the seed's grading contract",
                detail=seed.id,
            )
            continue

        kept.append(GeneratedQuestion(source_id=seed.id, text=paraphrase.text))
        tracer.record(kind="code", decided_by="code", title="Keep: new and verified", detail=seed.id)

    out_path.parent.mkdir(parents=True, exist_ok=True)
    lines = [json.dumps({"source_id": g.source_id, "question": g.text}, sort_keys=True) for g in kept]
    out_path.write_text("\n".join(lines) + ("\n" if lines else ""), encoding="utf-8", newline="\n")
    tracer.record(kind="code", decided_by="code", title="Write verified synthetic questions", detail=out_path.name)

    return SyntheticResult(kept=kept, rejected=rejected, out_path=out_path)
```

Nothing here is a model decision: the code generates, checks, asks and grades in the same order
every time, so every recorded step is `decided_by: "code"`. Against the real set, `python -m
examples.synthetic_data --model stub:scripted --out .local/scratch/synthetic-data/questions.jsonl`
keeps 28 and rejects 4, three for repeating text already seen and one at verification, with the
counts coming from paraphrases written down in advance rather than from a model. The same command
with `--model stub` keeps nothing and rejects all 32 at verification: the echoing stub's
placeholder answer matches no real grading pattern. That is the check working on an input that was
never trying to be right.

How much it actually catches is worth being precise about, so the tests attack it. A paraphrase
that changes a number ("with the top rack removed") is rejected, because the blind answer states
the new number and the seed's pattern wants the old one. A negation is rejected for the same
reason: *unless* the answer denies the seed's own fact in the seed's own words, and "It does not
hold 12 place settings" contains "12 place setting", so that one passes. That case is pinned as a
test rather than left to be discovered: this is a pattern check, not a meaning check, and a
paraphrase that quietly changed the question can still clear it. Diversity is the other blind
spot. Thirty-two paraphrases that differ from each other and reword every question the same way
grammatically pass everything here; the five question kinds in `docs/EVALS.md` (lookup,
multi-hop, numeric, unanswerable, conflicting sources) are the structural variety a real
generated set has to be checked for on top of text-level deduplication.

Argilla's distilabel does this at a scale a 50-line example does not: its README calls it "the
framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable
pipelines based on verified research papers"[3], and the same README opens by saying "The
original authors have moved on to other projects" and that community collaborators have joined to
maintain it[3]: worth knowing before a pipeline depends on it.

## When you do not need this

Count what you already have before generating anything. A handful of real examples that cover the
task, or an afternoon of someone writing the missing ones, beats a generated set outright: real
examples need no filter before you can trust them, and building a filter you can trust is most of
this work.

Two conditions have to hold together before generating is the cheaper road. Real examples must be
genuinely too rare, too expensive or too slow to collect at the volume a training or eval set needs:
the situation Self-Instruct's paper set out to make cheaper[1]. And somebody must
actually be going to build the verification step, rather than skip it because generating was the
interesting half. One without the other produces volume, not data.

## Failure modes

### Generated data used with no filter at all

- **How to notice it:** A dataset is generated and trained on directly, with nothing checking whether any individual example is correct, diverse, or even different from another example already in the set.
- **How to test for it:** Ask what checked a sample of the generated set before it was used. 'A model wrote it' is not an answer to that question.

### A filter that only checks surface form

- **How to notice it:** The checks on this page catch a repeat and an answer that stopped matching the seed's patterns. What they cannot catch is a paraphrase whose meaning changed but whose blind answer still contains the seed's accept text: a negation is the easy case, since denying a fact repeats it.
- **How to test for it:** Hand-read a sample of what the filter kept, comparing each paraphrase's meaning against its seed question rather than its verdict. The example's own test suite pins one paraphrase that passes and should not.

### Model collapse from training on an unchecked chain

- **How to notice it:** A generated set is used to train a model, whose own output later becomes the seed for the next round of generation, with no checked, real data reentering the loop.
- **How to test for it:** Trace where each generation's seed data came from. If it is entirely the previous generation's own unchecked output, the loop the 2023 model-collapse paper describes is the one running.

### Narrow generation mistaken for broad coverage

- **How to notice it:** A large generated set looks comprehensive by its count, but every example was produced from the same handful of seed questions or the same prompt template, so it covers less variety than its size suggests.
- **How to test for it:** Check how many distinct seeds or templates the set was generated from, not just how many examples came out the other end.

### Leakage between a generated training set and the real eval set

- **How to notice it:** A paraphrase generated for training turns out to be close enough to a question already in the eval set that training on it inflates a later score on that same question.
- **How to test for it:** Check generated text against every question the eval set holds, not only against the seeds it was generated from. This is the wider net this page's example casts, and it still only catches identical text once normalized, never a genuine rewording.

## How to Evaluate It

The example produces questions rather than answers, so the grading contract in
`evals/questions.json` has nothing to grade it against (see `docs/EVALS.md`). What it reports
instead is yield and reasons: how many of the 32 seeds produced a kept paraphrase, and how many
fell to a duplicate versus a failed verification. Watch the ratio rather than the total. A run that
suddenly keeps more has usually loosened a check.

For a generated set meant to train or test something else, the measurement is downstream and
comparative: score the thing that was trained or tested on generated data the ordinary way, score
the same thing built from real data only, and report both. A yield figure from the generator says
nothing about either.

## Run it

**What to monitor.** The rejection breakdown (duplicate versus failed verification) on every generation run,
  not only the final kept count; a rejection rate that suddenly drops usually means the checks
  loosened, not that the generator improved.

**Cost at volume.** Two model calls per seed question here (paraphrase, then blind answer), so cost
  scales with how many candidates are generated, not how many are kept: a low yield after
  verification means paying for calls whose output gets thrown away, which is the price of
  checking before training on any of it.

**How it fails in production.** A verification check that was tuned for one seed set's shape (a fixed
  appliance-support format, say) silently waves through generated data from a different domain,
  because nothing about the check was specific to what made the original examples right.

**What to log.** The generator model id, the seed each example came from, and which check it passed
  or failed, so a bad example downstream traces back to whether it was a generation problem or a
  gap in the filter.

## Try it

1. **Use it.** Find a product or paper reporting a dataset size built with a model. Search its page for 'filter', 'verify' or 'check'. A count with nothing beside it says how much was generated and nothing about how much was good.
2. **Build it.** Run python -m examples.synthetic_data --model stub:scripted --out .local/scratch/synthetic-data/questions.jsonl from the repo root and read the two rejection reasons it prints. Then open examples/synthetic_data/run.py and change ANSWER_SYSTEM_PROMPT to ask for a one-word answer; against a real model, would the failed-verification count go up or down, and why?
3. **Either lane.** Take one question from evals/questions.json and write two paraphrases by hand: one that keeps the same answer, one that quietly changes what is asked while still sounding like a paraphrase. Check both against grade_exact with the original's contract. Does it catch the second?


## Sources

1. [Self-Instruct: Aligning Language Models with Self-Generated Instructions](https://arxiv.org/abs/2212.10560) — arXiv (University of Washington and others), 2022-12-20 (accessed 2026-09-19)
2. [The Curse of Recursion: Training on Generated Data Makes Models Forget](https://arxiv.org/abs/2305.17493) — arXiv (Shumailov, Shumaylov, Zhao, Gal, Papernot, Anderson), 2023-05-27 (accessed 2026-09-19)
3. [distilabel](https://github.com/argilla-io/distilabel) — Argilla (accessed 2026-09-19)


Last reviewed 2026-09-19.
