Topics at every level

Changing the model

Fine-tuning, distillation, synthetic data and automated prompt tuning.

Sourced

Concept at a glance

Choose what you want to change.

Decision pathsConceptual illustration
Choose what you want to change.Observed weakness leads to Model weights. Observed weakness leads to Training data. Observed weakness leads to Prompt. Changing weights, training examples, and prompts are different interventions.Observed weaknessFind what needs improvementModel weightsFine-tune or distillTraining dataGenerate and verify examplesPromptSearch for betterinstructionsChoose what you want to change.Observed weakness leads to Model weights. Observed weakness leads to Training data. Observed weakness leads to Prompt. Changing weights, training examples, and prompts are different interventions.Observed weaknessFind what needs improvementModel weightsFine-tune or distillTraining dataGenerate and verify examplesPromptSearch for betterinstructions
Read the connections in words
  • Observed weakness → Model weights: Fine-tune or distill.
  • Observed weakness → Training data: Generate and verify examples.
  • Observed weakness → Prompt: Search for better instructions.
Key idea

Changing weights, training examples, and prompts are different interventions.

A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Changing the model: see it in practice.

Choosing how to improve task performance through context, prompts, data, or changes to model weights.

What you’ll walk through

Follow a recurring model failure into a choice of improvement method. Compare changing instructions, supplying better information, and changing model behavior before committing to training.

The task in this version

Improve a classifier that confuses access and billing issues.

What you’ll learn to check

A baseline error set, intervention comparison, held-out evaluation plan, and a justified choice of the simplest adequate method.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Improve a classifier that confuses access and billing issues.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Audit: locked-invoice examples mislabeled. Prompt and label definitions disagree.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Different failures have different causes. Missing current facts, unclear labels, and inconsistent formatting should not automatically receive the same treatment.

1 / 6
In this topic

4 pages under changing the model

Each one goes further into a part of this page than this page does.

Fine-tuning and adapters

Sourced

Training a model further on your own examples, in full or with small adapters such as LoRA.

Distillation

Sourced

Training a smaller model to reproduce what a larger one does on your task.

Synthetic data

Sourced

Using a model to write training or test examples, and checking them before they are used.

Prompt optimization

Sourced

Letting a program search for better prompts against a test set.

Apply this to your project

Describe your task to your own model and use Changing the model as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Every other technique on this site changes what you send the model, on every call. Adaptation changes the model itself, once, so a later call can be shorter, cheaper or more consistent without repeating the same instructions or examples. Fine-tuning trains further on your own examples. LoRA and other adapters train a small add-on instead of the whole model: the LoRA paper describes freezing the pretrained weights and injecting trainable “rank decomposition matrices” into each layer, and reports that this cuts trainable parameters by 10,000 times and GPU memory by 3 times against fine-tuning GPT-3 175B with Adam[1]. Distillation trains a smaller model on a larger one’s outputs. Reinforcement fine-tuning trains against a scored reward instead of fixed example answers. Synthetic data generates training examples with a model. Automated prompt optimization searches for a better prompt instead of a person hand-editing one.

Adaptation is orthogonal to the eight levels: an adapted model can sit under a single chat call or under one seat in a team of agents. It is the alternative to writing the same instructions into every prompt: prompt engineering’s “say it every time” against adaptation’s “train it in once.”

This page is sourced, not measured: the choice below is described from primary sources, but no training run has happened here, so what follows shows how the data gets prepared and nothing more.

Practical guidance

The pitch to test is “smaller, cheaper, or trained on our own material, and it matches the one you use today.” Build a switching test before you believe it. Pull twenty real tasks from your own recent work (questions you actually got asked, drafts you actually wrote) and write down, for each, the answer you already know is right. Send all twenty to the tool you use now and to the one being pitched, then read the two sets of answers side by side, task by task, not score by score.

Two disagreements out of twenty is worth a closer look before you switch; five or more means the new tool is not ready to replace the old one on your actual work, whatever the pitch says. Read every disagreement rather than just counting them: a wrong answer on something you handle every week matters more than one on something you rarely hit.

Ask what the “trained on our docs” or “trained on our data” claim actually covers, the same capture-then-filter step behind most distilled or fine-tuned products: a maker runs a larger model over examples of one task, keeps what meets its bar, and trains a smaller model on that[3]. A model narrowed to one job that way can be excellent at that job and confidently wrong the moment you ask it something else. Your twenty-task file only catches that if a few of the twenty sit outside the narrow job the vendor is actually selling.

“Trained on our docs” is also not “reads our docs.” A model trained further on your material got more consistent at the kind of thing it saw during training; it did not gain a live lookup of that material. A fact from last week is not something training put there, so a question about something recent tests memory the tool does not have. If the product will not also let you attach the current document and answer from that, treat “it knows our docs” as a guess dressed as a fact.

Switching to something with no vendor at all, such as writing a longer, more specific prompt for the tool you already use, needs the same twenty-task comparison before you call it better, not just cheaper.

Implementation details

Nothing on this site trains a model: no example here calls a fine-tuning API, and no result file exists for any adaptation technique. What a builder can do without one is prepare the data a fine-tuning job would actually need, and check that it is not broken before spending anything on a training run.

The example turns the site’s own 60-question set into a small supervised fine-tuning file. Each question becomes one line in chat format (a system/user/assistant message list), the shape Together AI’s fine-tuning data preparation guide documents: each message has “a role (system, user, or assistant) and content”, and a conversation “must start with system or user and alternate user and assistant afterwards”[2]. The same guide recommends holding out a validation file rather than training on everything, so training progress can be checked against examples the run never saw[2]; the example splits the 60 questions into a training file and a validation file with a fixed random seed, so the same seed always produces the same split.

examples/adaptation/run.py · lines 61–66
def load_examples(questions_path: Path = DEFAULT_QUESTIONS_PATH) -> list[Example]:
    """Every question in the set, as a training example. Every question carries a plain-language
    `answer` even when its kind is `unanswerable` (the correct completion there is the model
    saying so), so nothing is filtered out by kind."""
    data = json.loads(questions_path.read_text(encoding="utf-8"))
    return [Example(id=q["id"], question=q["question"], answer=q["answer"]) for q in data["questions"]]

A held-out split only means something if training never saw the held-out questions under a different guise. leaked_questions checks every validation question’s normalized text against the training file’s, and reports any that show up in both:

examples/adaptation/run.py · lines 83–92
def leaked_questions(train: list[Example], val: list[Example]) -> list[str]:
    """Val-set ids whose normalized question text also appears in the train set. A held-out split
    is only worth anything if training never saw the same question under a different id.

    This catches only questions that are identical once normalized. A paraphrase in genuinely
    different words ("how long is the warranty" against "what is the warranty period") is a
    near-duplicate this equality test cannot see; catching those needs a similarity measure, and
    an embedding of each question is the usual one."""
    train_texts = {_normalize(ex.question) for ex in train}
    return sorted(ex.id for ex in val if _normalize(ex.question) in train_texts)

_normalize lowercases, drops punctuation and collapses whitespace, so two questions that differ only in styling count as one. On this set the check comes back empty: all 60 questions normalize to distinct strings. The test file proves it catches a leak by putting the same question, restyled, on opposite sides of the split, and proves what it misses, by letting a genuine paraphrase through. Equality over normalized text is the floor. Catching paraphrases needs a similarity measure instead, usually an embedding of each question, and that is the gap that starts to matter on a larger set built partly from generated questions: the job Argilla’s distilabel describes itself as built for, “a framework for synthetic data and AI feedback”[6].

Two techniques this page covers have no data-preparation step to show. Reinforcement fine-tuning trains against a grader’s score rather than fixed example answers: OpenAI’s own guide says the method “samples several responses per prompt, scores them with the grader, and applies policy-gradient updates based on those rewards”[4], so there is no training file to build at all. Automated prompt optimization tunes a prompt or its few-shot examples against a metric instead of retraining weights; DSPy’s own description is “algorithms for optimizing their prompts and weights” so a program does not depend on “brittle prompts”[5], which is closer to what this site’s own eval loop would drive than to a training file.

Which of these a maker currently sells changes faster than the methods do. Both OpenAI guides cited here carried this notice when they were read for this page: “OpenAI is winding down the fine-tuning platform. The platform is no longer accessible to new users, but existing users of the fine-tuning platform will be able to create training jobs for the coming months”[3][4]. Fine-tuned models stay available for inference until their base models are deprecated, the same notice adds; the reinforcement fine-tuning guide also limits that method to one reasoning model id[4]. Together AI’s fine-tuning documentation carries no such notice[2]. Read the maker’s own page before planning around a service; the methods outlast the platforms that sell them.

When you do not need this

Try a better prompt, a few-shot example, or RAG first. Most of what looks like a reason to fine-tune is actually a prompt problem (the instructions were not specific enough) or a retrieval problem (the model was never given the fact it needed), and both are cheaper to fix and faster to test than a training run.

Adaptation earns its cost once the same instructions are sent enough times that training them in once is cheaper than repeating them on every call, or once a task needs a model to behave more consistently than any prompt can reliably hold it to. Neither condition is about the model being wrong on one specific question, which prompting and retrieval already fix more cheaply; both are about the shape and volume of the calls.

Failure modes

A validation leak inflates the score

How to notice it
Validation accuracy looks strong but real traffic performs worse, because a near-duplicate of a validation question was also present, reworded, in the training file.
How to test for it
Run the example's leaked_questions check, or an embedding-similarity version of it, on the actual split before trusting a validation number; exact-text matching alone lets a reworded duplicate through.

Narrow training mistaken for broad knowledge

How to notice it
A distilled or fine-tuned model handles the task it was trained for well, then confidently gets something outside that task wrong in a way the larger model it was trained from would not have.
How to test for it
Ask the adapted model a question clearly outside the narrow task it was trained for and compare the answer against the base model's; a gap that only appears outside the training task is this failure.

Trained facts read as current facts

How to notice it
The model states something it learned during training as fact, with nothing in the answer flagging that the world may have moved on since the training data was collected.
How to test for it
Ask about something in the training domain that has since changed, with no document attached, and check whether the model states the old fact with the same confidence as a current one.

The training platform is wound down

How to notice it
A fine-tuning or reinforcement fine-tuning job that used to work can no longer be created, though inference on models already trained keeps working, because the maker retired the training service without retiring what it produced.
How to test for it
Read the maker's own guide for a notice like the one this page quotes before planning around a training service, not just the date the last job was submitted.

A reward the grader can game

How to notice it
A reinforcement-fine-tuned model's score against its own reward model climbs while its answers, read by a person, do not actually improve. This is the same grader-hacking risk this site's evals topic covers, applied to training instead of testing.
How to test for it
Hand-check a sample of the reward grader's own verdicts the way this site's eval runner checks a rubric grader's, rather than trusting the trend of the reward curve alone.

At each level

  • Conventional software: a classical model such as the ones on the level 0 page is, by definition, already fit to your data every time it is trained: the questions this page raises about adaptation do not arise until there is a language model to adapt.
  • Direct prompting: a fine-tuned model can replace a long, repeated prompt-engineered system prompt with a shorter call that already behaves the trained way, at the cost of retraining whenever the instructions change.
  • Added context: adaptation does not substitute for retrieval: training a model further changes how it behaves, not what current facts it can reliably recall, so a fine-tuned model still needs documents in front of it for anything that changes after training.
  • Workflows: a fixed step that is called the same way thousands of times, with the same instructions and shape of input every time (one step of prompt chaining, say) is the cheapest place to swap a small adapted model in for a large general one.
  • Tool use: a model fine-tuned on your own tool set can produce fewer malformed function calls than a general model prompted with the same tool definitions, since it has seen your schemas specifically rather than schemas in general.
  • Agent loops: reinforcement fine-tuning fits a single agent’s loop especially well, because it trains directly against whether the task got done, the same thing the loop’s own stop decision is trying to get right, instead of imitating example transcripts.
  • Teams of Agents: a fixed role played by one agent every time (an author, a reviewer) is a narrow, repeated task, the case adaptation is built for; each seat could run its own smaller adapted model instead of every seat running the same large one.
  • Always-on agents: an always-on assistant logs every action it proposed and what happened to it, which is exactly the raw material a later fine-tuning or distillation pass would use, and exactly the data the safety, privacy and governance topic’s questions about retention and training use apply to.

Practices

  • Before fine-tuning, try the cheaper alternative first: a better prompt, a few-shot example, or retrieval. Adaptation is worth the cost when the same instructions are sent enough times that training once is cheaper than prompting every time, not by default.
  • Hold out a validation split and never train on it. A model that has only ever been graded on data it was also trained on has not been graded.
  • Check a prepared training file for near-duplicate examples straddling the train/validation split before spending anything on a training run; a leak inflates validation scores without improving anything real.
  • Do not treat a fine-tuned model as a substitute for giving it current documents. Ask what the model was trained to do, separately from what it was trained to know.
  • Keep the data a fine-tuning job is built from somewhere it can be regenerated or re-audited, the same discipline evals/questions.json follows for this site’s own eval set.
  • Four pages under this one go further than this page does on each way of adapting a model: fine-tuning and adapters, distillation, synthetic data and prompt optimization.

Run it

What to monitor

Whether the fine-tuned model's behavior on real traffic still matches what the validation split predicted; a gap that grows over time usually means real inputs have drifted away from the training data's shape, not that the model got worse.

Cost at volume

A training run is a one-time or periodic cost; a fine-tuned or distilled small model can then cost less per call than a large general model prompted the long way, but only on the narrow task it was trained for: traffic outside that task still needs the general model.

How it fails in production

The world changes and the training data doesn't: a model fine-tuned on last quarter's product line answers confidently and wrong about this quarter's, with nothing in its own output flagging that its training predates the change.

What to log

The training data's version or commit, the base model id, and the date of the run, so a later question about why the model behaves a certain way can be traced to what it was actually trained on.

Try it

  1. Use it

    Find a product that advertises a "custom-trained" or "fine-tuned" small model. Read what task it claims to be good at, and ask whether that claim is about behavior on that one task or about general knowledge: the page usually only supports the first.

  2. Build it

    Run python -m examples.adaptation --out .local/scratch/adaptation from the repo root, open val.jsonl, and check that every line has exactly one system, one user and one assistant message in that order.

  3. Either lane

    Pick two questions from evals/questions.json that ask about the same appliance in different words, and decide whether they're similar enough that putting one in training and one in validation would leak. Write down what made the call.

How it connects

Before, after and instead of this

Instead of

Optional: products, tools, and models

9 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 3 more examples
  • Transformers Tool or framework · Hugging Face

    Model library

    Checked 09/18/2026
  • TRL Tool or framework · Hugging Face

    Fine-tuning library

    Checked 09/18/2026
  • Unsloth Tool or framework · Unsloth

    Fine-tuning library

    Checked 09/18/2026
In practice

Improve extraction for your documents

Identify the failure first, then test whether better examples, a revised prompt, or model training addresses it.

Out there

Named products, tools and models

Tools10
  • Axolotlopen source · fine-tuning library
  • distilabelArgilla · synthetic data
  • DSPyStanford NLP · prompt programs and optimizers
  • MLXApple · training and inference on Apple hardware
  • OpenAI fine-tuningOpenAI · hosted fine-tuningRetired 2026
  • PEFTHugging Face · LoRA and other adapters
  • Together AI fine-tuningTogether AI · hosted fine-tuning
  • TransformersHugging Face · model library
  • TRLHugging Face · fine-tuning library
  • UnslothUnsloth · fine-tuning library

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. LoRA: Low-Rank Adaptation of Large Language Models · arXiv (Microsoft), 06/17/2021 (accessed 09/19/2026)
  2. Data preparation · Together AI (accessed 09/19/2026)
  3. Supervised fine-tuning · OpenAI (API documentation) (accessed 09/19/2026)
  4. Reinforcement fine-tuning · OpenAI (API documentation) (accessed 09/19/2026)
  5. DSPy · Stanford NLP (accessed 09/19/2026)
  6. distilabel · Argilla (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page