# Fine-tuning and adapters

_Topics at every level · sourced_

Training a model further on your own examples, in full or with small adapters such as LoRA.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a stable behavior requirement through training-data preparation and evaluation of a candidate model. Inspect label quality and generalization rather than assuming training guarantees learning the intended rule.

**Assumptions:** Training examples must represent the target behavior and be appropriate to use. A model can learn annotation inconsistencies or irrelevant cues.

**Design choices:** Compare with prompting and retrieval first. Separate training, development, and final evaluation data, and keep a baseline that has not been adapted.

**Request:** Plan training for our stable support taxonomy.

**Starting evidence:** Labels: access, billing, hardware. Repeated messages from the same incidents appear in the dataset.

**Action and control:** Audit labels and split by incident to prevent near-duplicate leakage. No training runs here.

**Stage records (authored, not executed):**

### Input record

Labels: access, billing, hardware. Repeated messages from the same incidents appear in the dataset.

What changed: Establish the facts supplied for this version of the task.

### Design note

Compare with prompting and retrieval first. Separate training, development, and final evaluation data, and keep a baseline that has not been adapted.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Audit labels and split by incident to prevent near-duplicate leakage. No training runs here.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Plan: clean training set, development selection, untouched incident-separated test set. No accuracy gain claimed.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Dataset split, label audit, a clearly illustrative training artifact, and held-out before/after results only when real measurements exist.

If the result falls short:
If a class improves while others regress, examine label definitions and data balance. Retain the previous model until the tradeoff is acceptable for the task.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to repeated, stable behavior patterns. Changing facts usually need maintained information sources; they are not automatically a reason to retrain.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Plan: clean training set, development selection, untouched incident-separated test set. No accuracy gain claimed.

**Change something — Put the same incidents in train and test:** A high score can reflect leakage. Rebuild the split before comparing models.

**Decision:** Does accuracy with leaked duplicates prove generalization?

**Answer:** No; use an independent split.

**Why:** Prevent train/test leakage and preserve rare categories; training is not a reliable store for frequently changing facts.

**Review criteria:** Dataset split, label audit, a clearly illustrative training artifact, and held-out before/after results only when real measurements exist.

**Recovery:** If a class improves while others regress, examine label definitions and data balance. Retain the previous model until the tradeoff is acceptable for the task.

**Adapt it:** Apply this to repeated, stable behavior patterns. Changing facts usually need maintained information sources; they are not automatically a reason to retrain.

Fine-tuning trains a model further on your own examples. It belongs to
[changing the model](/gradient_ascent/techniques/adaptation/), a topic that runs across the eight
levels instead of sitting on one: a fine-tuned model can answer a single chat call or fill one seat
in a team of agents.

It comes in two shapes. Full fine-tuning moves every weight, at a higher compute and memory cost.
An adapter leaves the original weights alone and trains a small add-on instead: the LoRA paper
describes freezing the pretrained weights and injecting trainable "rank decomposition matrices"
into each layer of the Transformer architecture, and reports, "Compared to GPT-3 175B
fine-tuned with Adam, LoRA can reduce the number of trainable parameters by 10,000 times and the
GPU memory requirement by 3 times."[1] Those are two figures measured against that one
model trained with that one optimizer, not a multiplier to expect from any job. Together AI calls
LoRA "the default training mode" on its own hosted service[2].

Two methods covered below train against something other than a fixed correct answer: preference
tuning ranks one response over another, and reinforcement fine-tuning trains against a scored
reward. The siblings under this topic are
[distillation](/gradient_ascent/techniques/distillation/),
[synthetic data](/gradient_ascent/techniques/synthetic-data/) and
[prompt optimization](/gradient_ascent/techniques/prompt-optimization/).

This page is sourced, not measured: what fine-tuning costs and changes comes from the makers' own
documentation, and no training run has happened here.

## Practical guidance

Someone offering to fine-tune a model for your team is asking you to pay for a training run now in
exchange for shorter, cheaper, more consistent answers later. Two questions, put to the vendor in
writing, decide whether that is worth it, and one platform's own numbers give you a floor to hold
their answers against.

How much of your own data does the job actually need? OpenAI's current guide states, "The minimum
number of examples you can provide for fine-tuning is 10." It adds, "We see improvements from
fine-tuning on 50–100 examples, but the right number for you varies greatly and depends on the use
case," and recommends "starting with 50 well-crafted demonstrations and evaluating the results."
Past that point its own advice is not more data: "If 50 examples have no impact, rethink your task
or prompt before adding training data"[5]. Those are OpenAI's numbers for OpenAI's
platform, not a law of fine-tuning anywhere else, but a vendor asking for thousands of examples
before they will even start owes you a reason theirs needs so many more.

Will the training service still be open when you need to retrain? The same documentation states,
"OpenAI is winding down the fine-tuning platform. The platform is no longer accessible to new
users, but existing users of the fine-tuning platform will be able to create training jobs for the
coming months."[4] It adds, "All fine-tuned models will remain available for inference
until their base models are deprecated." So a model already trained keeps answering after that.
Together AI's fine-tuning overview carries no such notice today[2]. Ask any vendor
selling training directly whether the service is still taking new jobs, and for how long.

The check that actually tells you whether it worked: collect real questions you already know the
right answer to, run them through the fine-tuned model and the one it replaces, and read both sets
of answers side by side rather than trusting either party's summary number.

And it will not teach the model something that happened last week. Fine-tuning changes weights
once, in advance; it does not give the model a live lookup of a document, so a question about
something newer than the training data gets a guess, not a fact.

## Implementation details

Nothing on this site trains a model: no example here calls a fine-tuning API, on this page or on
[the adaptation page](/gradient_ascent/techniques/adaptation/) it hangs from. What
`examples/adaptation` builds instead is the file a supervised fine-tuning job would actually need:
a chat-format JSONL split into training and validation questions from the site's own 60-question
set, with a check that no question leaks across the split. The adaptation page shows the part that
turns a question into a training example and the leak check itself; this page shows the split those
two functions sit between.

`split` shuffles with a fixed seed and cuts a validation fraction off the top, so the same seed
always produces the same partition:

`examples/adaptation/run.py` (lines 95-99)

```python
def split(examples: list[Example], *, val_fraction: float, seed: int) -> tuple[list[Example], list[Example]]:
    shuffled = list(examples)
    random.Random(seed).shuffle(shuffled)
    val_count = max(1, round(len(shuffled) * val_fraction))
    return shuffled[val_count:], shuffled[:val_count]  # train, val
```

The default, 20%, exists for the same reason Together AI's own fine-tuning data preparation guide
gives. Its instructions for carving "a validation set out of a single JSONL file" continue:
"Then pass both files to the job and set `n_evals` above 0:" and, further on, "The model evaluates
against the validation set at the specified intervals" during training[3]. A held-out set
is only useful if something is scored against it while the job runs, not just kept aside.

`run` is the whole pipeline in order: load the questions, split them, write both files, then check
the split for a leak.

`examples/adaptation/run.py` (lines 108-139)

```python
def run(
    tracer: Tracer,
    *,
    questions_path: Path = DEFAULT_QUESTIONS_PATH,
    out_dir: Path,
    val_fraction: float = 0.2,
    seed: int = 0,
) -> BuildResult:
    examples = load_examples(questions_path)
    tracer.record(kind="code", decided_by="code", title="Load questions as chat examples", detail=f"{len(examples)} examples")

    train, val = split(examples, val_fraction=val_fraction, seed=seed)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Shuffle and split into train and validation",
        detail=f"{len(train)} train, {len(val)} val, seed={seed}",
    )

    train_path, val_path = out_dir / "train.jsonl", out_dir / "val.jsonl"
    _write_jsonl(train_path, train)
    _write_jsonl(val_path, val)
    tracer.record(kind="code", decided_by="code", title="Write JSONL files", detail=f"{train_path.name}, {val_path.name}")

    leaked = leaked_questions(train, val)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check for leaked questions between splits",
        detail=", ".join(leaked) or "none",
    )
    return BuildResult(train=train, val=val, train_path=train_path, val_path=val_path, leaked=leaked)
```

Every step it records is `decided_by: "code"`: the split is a fixed shuffle-and-cut, not a choice
a model makes, so this example contributes zero model-decided steps, the same as any other data
preparation step.

Preference tuning and reinforcement fine-tuning have no file to build the way supervised
fine-tuning does, because neither trains against one fixed right answer. Together AI's own
overview separates two of its training methods on exactly that line: supervised fine-tuning trains
"on demonstration data with one target completion per example," while preference fine-tuning is
described as, "Align a model with rankings over preferred and dispreferred responses using
DPO."[2] OpenAI documents its own DPO option in the same shape, telling a reader to
"Provide both a correct and incorrect example response for a prompt. Indicate the correct response
to help the model perform better."[4] It lists three model ids the method is available
for today: `gpt-4.1-2025-04-14`, `gpt-4.1-mini-2025-04-14` and `gpt-4.1-nano-2025-04-14`.
Reinforcement fine-tuning goes further still: OpenAI's own guide says that during training the
platform "samples several responses per prompt, scores them with the grader, and applies
policy-gradient updates based on those rewards."[6] The method "is supported on o-series
reasoning models only, and currently only for o4-mini"[6]. Both are OpenAI's terms for
OpenAI's platform, read September 19, 2026; another maker's DPO offering is its own to describe. A
training file for either would be pairs or grader code, not the chat JSONL this example writes.

Self-run alternatives exist for anyone who would rather not depend on a hosted platform at all:
Hugging Face's Transformers, TRL and PEFT, Unsloth, Axolotl and Apple's MLX all train adapters or
full weights on your own hardware, at the cost of running the training yourself.

## When you do not need this

Most teams that ask about fine-tuning do not need it, and three questions usually settle that
without spending anything. Can you write down what the model is getting wrong as a rule? Then it
is a prompt, and a few-shot example in that prompt is a same-afternoon test. Is it getting a
*fact* wrong? Then it was never given the fact, which is
[retrieval](/gradient_ascent/techniques/rag/)'s job and not a training one. Is it getting the
wrong one of several jobs? Then the fix is [routing](/gradient_ascent/techniques/routing/)
between prompts, not one model taught to do all of them.

What is left after those three is the case fine-tuning is actually for: a behavior you can
demonstrate but not describe, on a call that runs often enough that carrying the instructions in
every prompt costs more than training them in once. A handful of calls a day is not that case, and
neither is a failure that happened twice.

## Failure modes

### Not enough data to move the needle

- **How to notice it:** The fine-tuned model behaves the same as the base model on the task it was trained for, because the training set was too small or too repetitive to teach it anything the prompt didn't already say.
- **How to test for it:** Follow OpenAI's own test for this, applied to any platform: add examples in batches and re-evaluate; if fifty good examples changed nothing, the fix is the task or the prompt, not more data.

### A validation leak inflates the score

- **How to notice it:** Validation performance looks strong but real traffic is worse, because a near-duplicate of a validation question was also present, reworded, in the training file.
- **How to test for it:** Run the leak check shown on this page and the adaptation page against the actual split before trusting a validation number; it catches an exact or punctuation-only duplicate, not a genuine paraphrase.

### Catastrophic forgetting on the rest of the model

- **How to notice it:** A model fine-tuned hard on one task gets measurably worse at things it used to do fine, because training changed weights that were doing useful work outside the trained task, not only inside it.
- **How to test for it:** Before and after training, run the same handful of prompts from outside the trained task and compare the answers, not just the trained task's own score.

### The hosted platform stops taking new jobs

- **How to notice it:** A workflow built around retraining periodically can no longer submit a new job, though models already trained keep serving inference, because the maker wound the training service down without retiring what it produced.
- **How to test for it:** Read the maker's own current guide for a notice like the one this page quotes before planning around a training service, not just the date the last job succeeded.

### A reward the grader can game

- **How to notice it:** A reinforcement-fine-tuned model's score against its own grader climbs while answers read by a person do not improve. This is the training-time version of the grader-hacking risk the site's evals topic covers for testing.
- **How to test for it:** Hand-check a sample of the grader's own verdicts on the training data, the way a rubric grader's verdicts are hand-checked at eval time, rather than trusting the reward curve alone.

## How to Evaluate It

This site's own 60-question set does not score fine-tuning: the example on this page builds and
checks a training file, and answers no question about the corpus, so the grading contract in
`evals/questions.json` has nothing to check it against (see `docs/EVALS.md`). What a real
fine-tuning job would still want measured is the same question set run twice (once on the base
model, once on the fine-tuned one) so a claim of improvement is a before/after score on the same
60 questions rather than a training-time number alone, the way the [evals](/gradient_ascent/techniques/evals/) page's own rule requires. Alongside that: the leak check
this page shows, run against the real split before spending anything on a job, and a handful of
off-task prompts checked before and after, to catch the forgetting failure mode above.

## Run it

**What to monitor.** Whether the fine-tuned model's behavior on real traffic still matches the last
  validation run; a growing gap usually means real inputs have drifted from the training data's
  shape, not that the model got worse on its own.

**Cost at volume.** Training is a one-time or periodic cost; the fine-tuned model can then cost less
  per call than a large general model prompted the long way, but only on the narrow task it was
  trained for. Traffic outside that task still needs the general model.

**How it fails in production.** The world moves and the training data doesn't. A model trained on last
  quarter's catalog answers confidently and wrong about this quarter's, with nothing in its own
  output flagging that its training predates the change.

**What to log.** The training data's version, the base model id, the method (supervised, DPO,
  reinforcement) and the date of the run, so a question about the model's behavior later can be
  traced to what it was actually trained on.

## Try it

1. **Use it.** Find a product that advertises a small, fast, task-specific model. Look for whether its own page says LoRA/adapter, full fine-tuning, or doesn't say, and whether that silence changes how much you'd trust a claim that it 'matches' a bigger model.
2. **Build it.** Run python -m examples.adaptation --out .local/scratch/fine-tuning from the repo root, then open val.jsonl and count the lines. Change --val-fraction to 0.1 and run it again; does the count change the way split's docstring says it should?
3. **Either lane.** Pick two questions from evals/questions.json about the same appliance and decide whether they're close enough that training on one and validating on the other would leak. Run the leak check and see whether your judgment matches leaked_questions's.


## Sources

1. [LoRA: Low-Rank Adaptation of Large Language Models](https://arxiv.org/abs/2106.09685) — arXiv (Microsoft), 2021-06-17 (accessed 2026-09-19)
2. [Fine-tuning: overview](https://docs.together.ai/docs/fine-tuning/overview) — Together AI (accessed 2026-09-19)
3. [Fine-tuning: data preparation](https://docs.together.ai/docs/fine-tuning/data-preparation) — Together AI (accessed 2026-09-19)
4. [Model optimization](https://developers.openai.com/api/docs/guides/model-optimization) — OpenAI (API documentation) (accessed 2026-09-19)
5. [Supervised fine-tuning](https://developers.openai.com/api/docs/guides/supervised-fine-tuning) — OpenAI (API documentation) (accessed 2026-09-19)
6. [Reinforcement fine-tuning](https://developers.openai.com/api/docs/guides/reinforcement-fine-tuning) — OpenAI (API documentation) (accessed 2026-09-19)


Last reviewed 2026-09-19.
