# Changing the model

_Topics at every level · sourced_

Fine-tuning, distillation, synthetic data and automated prompt tuning.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a recurring model failure into a choice of improvement method. Compare changing instructions, supplying better information, and changing model behavior before committing to training.

**Assumptions:** Different failures have different causes. Missing current facts, unclear labels, and inconsistent formatting should not automatically receive the same treatment.

**Design choices:** Try the simplest intervention matched to the failure and compare on representative cases. Training is a candidate when the desired behavior is stable and adequate examples exist.

**Request:** Improve a classifier that confuses access and billing issues.

**Starting evidence:** Audit: locked-invoice examples mislabeled. Prompt and label definitions disagree.

**Action and control:** Fix taxonomy and prompt ambiguity before deciding to train; separate context changes from weight changes.

**Stage records (authored, not executed):**

### Input record

Audit: locked-invoice examples mislabeled. Prompt and label definitions disagree.

What changed: Establish the facts supplied for this version of the task.

### Design note

Try the simplest intervention matched to the failure and compare on representative cases. Training is a candidate when the desired behavior is stable and adequate examples exist.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Fix taxonomy and prompt ambiguity before deciding to train; separate context changes from weight changes.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

First intervention: clarify labels and audit examples. Reserve held-out cases before model comparison.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

A baseline error set, intervention comparison, held-out evaluation plan, and a justified choice of the simplest adequate method.

If the result falls short:
If the improvement does not transfer to held-out cases, revisit the diagnosis. More examples or a larger model may not address the underlying issue.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this decision process for classification, writing style, extraction, or assistance. Define the observed failure first and choose the intervention second.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** First intervention: clarify labels and audit examples. Reserve held-out cases before model comparison.

**Change something — Errors concern changing account facts:** Supply current authorized facts through context or retrieval; training is not a dependable live account store.

**Decision:** Should every quality problem lead to fine-tuning?

**Answer:** No; diagnose and compare simpler fixes.

**Why:** Start with error analysis; distinguish retrieval/prompt changes from actual weight training and avoid assuming training is necessary.

**Review criteria:** A baseline error set, intervention comparison, held-out evaluation plan, and a justified choice of the simplest adequate method.

**Recovery:** If the improvement does not transfer to held-out cases, revisit the diagnosis. More examples or a larger model may not address the underlying issue.

**Adapt it:** Use this decision process for classification, writing style, extraction, or assistance. Define the observed failure first and choose the intervention second.

Every other technique on this site changes what you send the model, on every call. Adaptation
changes the model itself, once, so a later call can be shorter, cheaper or more consistent
without repeating the same instructions or examples. Fine-tuning trains further on your own
examples. LoRA and other adapters train a small add-on instead of the whole model: the LoRA paper
describes freezing the pretrained weights and injecting trainable "rank decomposition matrices"
into each layer, and reports that this cuts trainable parameters by 10,000 times and GPU memory
by 3 times against fine-tuning GPT-3 175B with Adam[1]. Distillation trains a smaller
model on a larger one's outputs. Reinforcement fine-tuning trains against a scored reward instead
of fixed example answers. Synthetic data generates training examples with a model. Automated
prompt optimization searches for a better prompt instead of a person hand-editing one.

Adaptation is orthogonal to the eight levels: an adapted model can sit under a single chat call
or under one seat in a team of agents. It is the alternative to writing the same instructions
into every prompt: [prompt engineering](/gradient_ascent/techniques/prompt-engineering/)'s
"say it every time" against adaptation's "train it in once."

This page is sourced, not measured: the choice below is described from primary sources, but no
training run has happened here, so what follows shows how the data gets prepared and nothing more.

## Practical guidance

The pitch to test is "smaller, cheaper, or trained on our own material, and it matches the one you
use today." Build a switching test before you believe it. Pull twenty real tasks from your own
recent work (questions you actually got asked, drafts you actually wrote) and write down, for
each, the answer you already know is right. Send all twenty to the tool you use now and to the one
being pitched, then read the two sets of answers side by side, task by task, not score by score.

Two disagreements out of twenty is worth a closer look before you switch; five or more means the
new tool is not ready to replace the old one on your actual work, whatever the pitch says. Read
every disagreement rather than just counting them: a wrong answer on something you handle every
week matters more than one on something you rarely hit.

Ask what the "trained on our docs" or "trained on our data" claim actually covers, the same
capture-then-filter step behind most distilled or fine-tuned products: a maker runs a larger model
over examples of one task, keeps what meets its bar, and trains a smaller model on
that[3]. A model narrowed to one job that way can be excellent at that job and
confidently wrong the moment you ask it something else. Your twenty-task file only catches that
if a few of the twenty sit outside the narrow job the vendor is actually selling.

"Trained on our docs" is also not "reads our docs." A model trained further on your material got
more consistent at the kind of thing it saw during training; it did not gain a live lookup of that
material. A fact from last week is not something training put there, so a question about something
recent tests memory the tool does not have. If the product will not also let you attach the
current document and answer from that, treat "it knows our docs" as a guess dressed as a fact.

Switching to something with no vendor at all, such as writing a longer, more specific prompt for
the tool you already use, needs the same twenty-task comparison before you call it better, not
just cheaper.

## Implementation details

Nothing on this site trains a model: no example here calls a fine-tuning API, and no result file
exists for any adaptation technique. What a builder can do without one is prepare the data a
fine-tuning job would actually need, and check that it is not broken before spending anything on
a training run.

The example turns the site's own 60-question set into a small supervised fine-tuning file. Each
question becomes one line in chat format (a `system`/`user`/`assistant` message list), the shape
Together AI's fine-tuning data preparation guide documents: each message has "a role (`system`,
`user`, or `assistant`) and `content`", and a conversation "must start with `system` or `user`
and alternate `user` and `assistant` afterwards"[2]. The same guide recommends holding
out a validation file rather than training on everything, so training progress can be checked
against examples the run never saw[2]; the example splits the 60 questions into a
training file and a validation file with a fixed random seed, so the same seed always produces
the same split.

`examples/adaptation/run.py` (lines 61-66)

```python
def load_examples(questions_path: Path = DEFAULT_QUESTIONS_PATH) -> list[Example]:
    """Every question in the set, as a training example. Every question carries a plain-language
    `answer` even when its kind is `unanswerable` (the correct completion there is the model
    saying so), so nothing is filtered out by kind."""
    data = json.loads(questions_path.read_text(encoding="utf-8"))
    return [Example(id=q["id"], question=q["question"], answer=q["answer"]) for q in data["questions"]]
```

A held-out split only means something if training never saw the held-out questions under a
different guise. `leaked_questions` checks every validation question's normalized text against
the training file's, and reports any that show up in both:

`examples/adaptation/run.py` (lines 83-92)

```python
def leaked_questions(train: list[Example], val: list[Example]) -> list[str]:
    """Val-set ids whose normalized question text also appears in the train set. A held-out split
    is only worth anything if training never saw the same question under a different id.

    This catches only questions that are identical once normalized. A paraphrase in genuinely
    different words ("how long is the warranty" against "what is the warranty period") is a
    near-duplicate this equality test cannot see; catching those needs a similarity measure, and
    an embedding of each question is the usual one."""
    train_texts = {_normalize(ex.question) for ex in train}
    return sorted(ex.id for ex in val if _normalize(ex.question) in train_texts)
```

`_normalize` lowercases, drops punctuation and collapses whitespace, so two questions that differ
only in styling count as one. On this set the check comes back empty: all 60 questions normalize
to distinct strings. The test file proves it catches a leak by putting the same question,
restyled, on opposite sides of the split, and proves what it misses, by letting a genuine
paraphrase through. Equality over normalized text is the floor. Catching paraphrases needs a
similarity measure instead, usually an embedding of each question, and that is the gap that
starts to matter on a larger set built partly from generated questions: the job Argilla's
distilabel describes itself as built for, "a framework for synthetic data and AI
feedback"[6].

Two techniques this page covers have no data-preparation step to show. Reinforcement fine-tuning
trains against a grader's score rather than fixed example answers: OpenAI's own guide says the
method "samples several responses per prompt, scores them with the grader, and applies
policy-gradient updates based on those rewards"[4], so there is no training file to
build at all. Automated prompt optimization tunes a prompt or its few-shot examples against a
metric instead of retraining weights; DSPy's own description is "algorithms for optimizing their
prompts and weights" so a program does not depend on "brittle prompts"[5], which is
closer to what this site's own eval loop would drive than to a training file.

Which of these a maker currently sells changes faster than the methods do. Both OpenAI guides
cited here carried this notice when they were read for this page: "OpenAI is winding down the
fine-tuning platform. The platform is no longer accessible to new users, but existing users of
the fine-tuning platform will be able to create training jobs for the coming
months"[3][4].
Fine-tuned models stay available for inference until their base models are deprecated, the same
notice adds; the reinforcement fine-tuning guide also limits that method
to one reasoning model id[4]. Together AI's fine-tuning documentation carries no such
notice[2]. Read the maker's own page before planning around a service; the methods
outlast the platforms that sell them.

## When you do not need this

Try a better prompt, a few-shot example, or [RAG](/gradient_ascent/techniques/rag/) first. Most
of what looks like a reason to fine-tune is actually a prompt problem (the instructions were not
specific enough) or a retrieval problem (the model was never given the fact it needed), and
both are cheaper to fix and faster to test than a training run.

Adaptation earns its cost once the same instructions are sent enough times that training them in
once is cheaper than repeating them on every call, or once a task needs a model to behave more
consistently than any prompt can reliably hold it to. Neither condition is about the model being
wrong on one specific question, which prompting and retrieval already fix more cheaply; both are
about the shape and volume of the calls.

## Failure modes

### A validation leak inflates the score

- **How to notice it:** Validation accuracy looks strong but real traffic performs worse, because a near-duplicate of a validation question was also present, reworded, in the training file.
- **How to test for it:** Run the example's leaked_questions check, or an embedding-similarity version of it, on the actual split before trusting a validation number; exact-text matching alone lets a reworded duplicate through.

### Narrow training mistaken for broad knowledge

- **How to notice it:** A distilled or fine-tuned model handles the task it was trained for well, then confidently gets something outside that task wrong in a way the larger model it was trained from would not have.
- **How to test for it:** Ask the adapted model a question clearly outside the narrow task it was trained for and compare the answer against the base model's; a gap that only appears outside the training task is this failure.

### Trained facts read as current facts

- **How to notice it:** The model states something it learned during training as fact, with nothing in the answer flagging that the world may have moved on since the training data was collected.
- **How to test for it:** Ask about something in the training domain that has since changed, with no document attached, and check whether the model states the old fact with the same confidence as a current one.

### The training platform is wound down

- **How to notice it:** A fine-tuning or reinforcement fine-tuning job that used to work can no longer be created, though inference on models already trained keeps working, because the maker retired the training service without retiring what it produced.
- **How to test for it:** Read the maker's own guide for a notice like the one this page quotes before planning around a training service, not just the date the last job was submitted.

### A reward the grader can game

- **How to notice it:** A reinforcement-fine-tuned model's score against its own reward model climbs while its answers, read by a person, do not actually improve. This is the same grader-hacking risk this site's evals topic covers, applied to training instead of testing.
- **How to test for it:** Hand-check a sample of the reward grader's own verdicts the way this site's eval runner checks a rubric grader's, rather than trusting the trend of the reward curve alone.

## At each level

- [Conventional software](/gradient_ascent/levels/0/): a classical model such as the ones on the
  [level 0 page](/gradient_ascent/techniques/order-zero/) is, by definition, already fit to your
  data every time it is trained: the questions this page raises about adaptation do not arise
  until there is a language model to adapt.
- [Direct prompting](/gradient_ascent/levels/1/): a fine-tuned model can replace a long, repeated
  [prompt-engineered](/gradient_ascent/techniques/prompt-engineering/) system prompt with a
  shorter call that already behaves the trained way, at the cost of retraining whenever the
  instructions change.
- [Added context](/gradient_ascent/levels/2/): adaptation does not substitute for
  [retrieval](/gradient_ascent/techniques/rag/): training a model further changes how it
  behaves, not what current facts it can reliably recall, so a fine-tuned model still needs
  documents in front of it for anything that changes after training.
- [Workflows](/gradient_ascent/levels/3/): a fixed step that is called the same way thousands of
  times, with the same instructions and shape of input every time (one step of
  [prompt chaining](/gradient_ascent/techniques/prompt-chaining/), say) is the cheapest place
  to swap a small adapted model in for a large general one.
- [Tool use](/gradient_ascent/levels/4/): a model fine-tuned on your own tool set can produce fewer
  malformed [function calls](/gradient_ascent/techniques/function-calling/) than a general model
  prompted with the same tool definitions, since it has seen your schemas specifically rather
  than schemas in general.
- [Agent loops](/gradient_ascent/levels/5/): reinforcement fine-tuning fits
  [a single agent](/gradient_ascent/techniques/single-agent/)'s loop especially well, because it
  trains directly against whether the task got done, the same thing the loop's own stop decision
  is trying to get right, instead of imitating example transcripts.
- [Teams of Agents](/gradient_ascent/levels/6/): a fixed role played by one agent every time (an
  author, a [reviewer](/gradient_ascent/techniques/debate-review/)) is a narrow, repeated task,
  the case adaptation is built for; each seat could run its own smaller adapted model instead of
  every seat running the same large one.
- [Always-on agents](/gradient_ascent/levels/7/): an
  [always-on assistant](/gradient_ascent/techniques/agent-teammates/) logs every action it
  proposed and what happened to it, which is exactly the raw material a later fine-tuning or
  distillation pass would use, and exactly the data the
  [safety, privacy and governance](/gradient_ascent/techniques/safety/) topic's questions about
  retention and training use apply to.

## Practices

- Before fine-tuning, try the cheaper alternative first: a better prompt, a few-shot example, or
  retrieval. Adaptation is worth the cost when the same instructions are sent enough times that
  training once is cheaper than prompting every time, not by default.
- Hold out a validation split and never train on it. A model that has only ever been graded on
  data it was also trained on has not been graded.
- Check a prepared training file for near-duplicate examples straddling the train/validation
  split before spending anything on a training run; a leak inflates validation scores without
  improving anything real.
- Do not treat a fine-tuned model as a substitute for giving it current documents. Ask what the
  model was trained to do, separately from what it was trained to know.
- Keep the data a fine-tuning job is built from somewhere it can be regenerated or re-audited,
  the same discipline `evals/questions.json` follows for this site's own eval set.
- Four pages under this one go further than this page does on each way of adapting a model:
  [fine-tuning and adapters](/gradient_ascent/techniques/fine-tuning/),
  [distillation](/gradient_ascent/techniques/distillation/),
  [synthetic data](/gradient_ascent/techniques/synthetic-data/) and
  [prompt optimization](/gradient_ascent/techniques/prompt-optimization/).

## Run it

**What to monitor.** Whether the fine-tuned model's behavior on real traffic still matches what the
  validation split predicted; a gap that grows over time usually means real inputs have drifted
  away from the training data's shape, not that the model got worse.

**Cost at volume.** A training run is a one-time or periodic cost; a fine-tuned or distilled small
  model can then cost less per call than a large general model prompted the long way, but only on
  the narrow task it was trained for: traffic outside that task still needs the general model.

**How it fails in production.** The world changes and the training data doesn't: a model fine-tuned on last
  quarter's product line answers confidently and wrong about this quarter's, with nothing in its
  own output flagging that its training predates the change.

**What to log.** The training data's version or commit, the base model id, and the date of the run,
  so a later question about why the model behaves a certain way can be traced to what it was
  actually trained on.

## Try it

1. **Use it.** Find a product that advertises a "custom-trained" or "fine-tuned" small model. Read what task it claims to be good at, and ask whether that claim is about behavior on that one task or about general knowledge: the page usually only supports the first.
2. **Build it.** Run python -m examples.adaptation --out .local/scratch/adaptation from the repo root, open val.jsonl, and check that every line has exactly one system, one user and one assistant message in that order.
3. **Either lane.** Pick two questions from evals/questions.json that ask about the same appliance in different words, and decide whether they're similar enough that putting one in training and one in validation would leak. Write down what made the call.


## Sources

1. [LoRA: Low-Rank Adaptation of Large Language Models](https://arxiv.org/abs/2106.09685) — arXiv (Microsoft), 2021-06-17 (accessed 2026-09-19)
2. [Data preparation](https://docs.together.ai/docs/fine-tuning/data-preparation) — Together AI (accessed 2026-09-19)
3. [Supervised fine-tuning](https://developers.openai.com/api/docs/guides/supervised-fine-tuning#distilling-from-a-larger-model) — OpenAI (API documentation) (accessed 2026-09-19)
4. [Reinforcement fine-tuning](https://developers.openai.com/api/docs/guides/reinforcement-fine-tuning) — OpenAI (API documentation) (accessed 2026-09-19)
5. [DSPy](https://github.com/stanfordnlp/dspy) — Stanford NLP (accessed 2026-09-19)
6. [distilabel](https://github.com/argilla-io/distilabel) — Argilla (accessed 2026-09-19)


Last reviewed 2026-09-19.
