Topics at every level

Fine-tuning and adapters

Training a model further on your own examples, in full or with small adapters such as LoRA.

Sourced

Concept at a glance

Learn from examples by changing model parameters.

SequenceConceptual illustration
Learn from examples by changing model parameters.Training examples leads to Training update. Training update leads to Adapted model. Training changes weights or adapters; it is different from adding context to one request.Training examplesInputs with desired outputsTraining updateWeights or small adaptersAdapted modelEvaluate on held-out tasksLearn from examples by changing model parameters.Training examples leads to Training update. Training update leads to Adapted model. Training changes weights or adapters; it is different from adding context to one request.Training examplesInputs with desired outputsTraining updateWeights or small adaptersAdapted modelEvaluate on held-out tasks
Read the connections in words
  • Training examples → Training update: Weights or small adapters.
  • Training update → Adapted model: Evaluate on held-out tasks.
Key idea

Training changes weights or adapters; it is different from adding context to one request.

A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Fine-tuning and adapters: see it in practice.

Further training model weights, fully or through adapters, on task-specific examples.

What you’ll walk through

Follow a stable behavior requirement through training-data preparation and evaluation of a candidate model. Inspect label quality and generalization rather than assuming training guarantees learning the intended rule.

The task in this version

Plan training for our stable support taxonomy.

What you’ll learn to check

Dataset split, label audit, a clearly illustrative training artifact, and held-out before/after results only when real measurements exist.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Plan training for our stable support taxonomy.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Labels: access, billing, hardware. Repeated messages from the same incidents appear in the dataset.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Training examples must represent the target behavior and be appropriate to use. A model can learn annotation inconsistencies or irrelevant cues.

1 / 6

Apply this to your project

Describe your task to your own model and use Fine-tuning and adapters as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Fine-tuning trains a model further on your own examples. It belongs to changing the model, a topic that runs across the eight levels instead of sitting on one: a fine-tuned model can answer a single chat call or fill one seat in a team of agents.

It comes in two shapes. Full fine-tuning moves every weight, at a higher compute and memory cost. An adapter leaves the original weights alone and trains a small add-on instead: the LoRA paper describes freezing the pretrained weights and injecting trainable “rank decomposition matrices” into each layer of the Transformer architecture, and reports, “Compared to GPT-3 175B fine-tuned with Adam, LoRA can reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times.”[1] Those are two figures measured against that one model trained with that one optimizer, not a multiplier to expect from any job. Together AI calls LoRA “the default training mode” on its own hosted service[2].

Two methods covered below train against something other than a fixed correct answer: preference tuning ranks one response over another, and reinforcement fine-tuning trains against a scored reward. The siblings under this topic are distillation, synthetic data and prompt optimization.

This page is sourced, not measured: what fine-tuning costs and changes comes from the makers’ own documentation, and no training run has happened here.

Practical guidance

Someone offering to fine-tune a model for your team is asking you to pay for a training run now in exchange for shorter, cheaper, more consistent answers later. Two questions, put to the vendor in writing, decide whether that is worth it, and one platform’s own numbers give you a floor to hold their answers against.

How much of your own data does the job actually need? OpenAI’s current guide states, “The minimum number of examples you can provide for fine-tuning is 10.” It adds, “We see improvements from fine-tuning on 50–100 examples, but the right number for you varies greatly and depends on the use case,” and recommends “starting with 50 well-crafted demonstrations and evaluating the results.” Past that point its own advice is not more data: “If 50 examples have no impact, rethink your task or prompt before adding training data”[5]. Those are OpenAI’s numbers for OpenAI’s platform, not a law of fine-tuning anywhere else, but a vendor asking for thousands of examples before they will even start owes you a reason theirs needs so many more.

Will the training service still be open when you need to retrain? The same documentation states, “OpenAI is winding down the fine-tuning platform. The platform is no longer accessible to new users, but existing users of the fine-tuning platform will be able to create training jobs for the coming months.”[4] It adds, “All fine-tuned models will remain available for inference until their base models are deprecated.” So a model already trained keeps answering after that. Together AI’s fine-tuning overview carries no such notice today[2]. Ask any vendor selling training directly whether the service is still taking new jobs, and for how long.

The check that actually tells you whether it worked: collect real questions you already know the right answer to, run them through the fine-tuned model and the one it replaces, and read both sets of answers side by side rather than trusting either party’s summary number.

And it will not teach the model something that happened last week. Fine-tuning changes weights once, in advance; it does not give the model a live lookup of a document, so a question about something newer than the training data gets a guess, not a fact.

Implementation details

Nothing on this site trains a model: no example here calls a fine-tuning API, on this page or on the adaptation page it hangs from. What examples/adaptation builds instead is the file a supervised fine-tuning job would actually need: a chat-format JSONL split into training and validation questions from the site’s own 60-question set, with a check that no question leaks across the split. The adaptation page shows the part that turns a question into a training example and the leak check itself; this page shows the split those two functions sit between.

split shuffles with a fixed seed and cuts a validation fraction off the top, so the same seed always produces the same partition:

examples/adaptation/run.py · lines 95–99
def split(examples: list[Example], *, val_fraction: float, seed: int) -> tuple[list[Example], list[Example]]:
    shuffled = list(examples)
    random.Random(seed).shuffle(shuffled)
    val_count = max(1, round(len(shuffled) * val_fraction))
    return shuffled[val_count:], shuffled[:val_count]  # train, val

The default, 20%, exists for the same reason Together AI’s own fine-tuning data preparation guide gives. Its instructions for carving “a validation set out of a single JSONL file” continue: “Then pass both files to the job and set n_evals above 0:” and, further on, “The model evaluates against the validation set at the specified intervals” during training[3]. A held-out set is only useful if something is scored against it while the job runs, not just kept aside.

run is the whole pipeline in order: load the questions, split them, write both files, then check the split for a leak.

examples/adaptation/run.py · lines 108–139
def run(
    tracer: Tracer,
    *,
    questions_path: Path = DEFAULT_QUESTIONS_PATH,
    out_dir: Path,
    val_fraction: float = 0.2,
    seed: int = 0,
) -> BuildResult:
    examples = load_examples(questions_path)
    tracer.record(kind="code", decided_by="code", title="Load questions as chat examples", detail=f"{len(examples)} examples")

    train, val = split(examples, val_fraction=val_fraction, seed=seed)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Shuffle and split into train and validation",
        detail=f"{len(train)} train, {len(val)} val, seed={seed}",
    )

    train_path, val_path = out_dir / "train.jsonl", out_dir / "val.jsonl"
    _write_jsonl(train_path, train)
    _write_jsonl(val_path, val)
    tracer.record(kind="code", decided_by="code", title="Write JSONL files", detail=f"{train_path.name}, {val_path.name}")

    leaked = leaked_questions(train, val)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check for leaked questions between splits",
        detail=", ".join(leaked) or "none",
    )
    return BuildResult(train=train, val=val, train_path=train_path, val_path=val_path, leaked=leaked)

Every step it records is decided_by: "code": the split is a fixed shuffle-and-cut, not a choice a model makes, so this example contributes zero model-decided steps, the same as any other data preparation step.

Preference tuning and reinforcement fine-tuning have no file to build the way supervised fine-tuning does, because neither trains against one fixed right answer. Together AI’s own overview separates two of its training methods on exactly that line: supervised fine-tuning trains “on demonstration data with one target completion per example,” while preference fine-tuning is described as, “Align a model with rankings over preferred and dispreferred responses using DPO.”[2] OpenAI documents its own DPO option in the same shape, telling a reader to “Provide both a correct and incorrect example response for a prompt. Indicate the correct response to help the model perform better.”[4] It lists three model ids the method is available for today: gpt-4.1-2025-04-14, gpt-4.1-mini-2025-04-14 and gpt-4.1-nano-2025-04-14. Reinforcement fine-tuning goes further still: OpenAI’s own guide says that during training the platform “samples several responses per prompt, scores them with the grader, and applies policy-gradient updates based on those rewards.”[6] The method “is supported on o-series reasoning models only, and currently only for o4-mini”[6]. Both are OpenAI’s terms for OpenAI’s platform, read September 19, 2026; another maker’s DPO offering is its own to describe. A training file for either would be pairs or grader code, not the chat JSONL this example writes.

Self-run alternatives exist for anyone who would rather not depend on a hosted platform at all: Hugging Face’s Transformers, TRL and PEFT, Unsloth, Axolotl and Apple’s MLX all train adapters or full weights on your own hardware, at the cost of running the training yourself.

When you do not need this

Most teams that ask about fine-tuning do not need it, and three questions usually settle that without spending anything. Can you write down what the model is getting wrong as a rule? Then it is a prompt, and a few-shot example in that prompt is a same-afternoon test. Is it getting a fact wrong? Then it was never given the fact, which is retrieval’s job and not a training one. Is it getting the wrong one of several jobs? Then the fix is routing between prompts, not one model taught to do all of them.

What is left after those three is the case fine-tuning is actually for: a behavior you can demonstrate but not describe, on a call that runs often enough that carrying the instructions in every prompt costs more than training them in once. A handful of calls a day is not that case, and neither is a failure that happened twice.

Failure modes

Not enough data to move the needle

How to notice it
The fine-tuned model behaves the same as the base model on the task it was trained for, because the training set was too small or too repetitive to teach it anything the prompt didn't already say.
How to test for it
Follow OpenAI's own test for this, applied to any platform: add examples in batches and re-evaluate; if fifty good examples changed nothing, the fix is the task or the prompt, not more data.

A validation leak inflates the score

How to notice it
Validation performance looks strong but real traffic is worse, because a near-duplicate of a validation question was also present, reworded, in the training file.
How to test for it
Run the leak check shown on this page and the adaptation page against the actual split before trusting a validation number; it catches an exact or punctuation-only duplicate, not a genuine paraphrase.

Catastrophic forgetting on the rest of the model

How to notice it
A model fine-tuned hard on one task gets measurably worse at things it used to do fine, because training changed weights that were doing useful work outside the trained task, not only inside it.
How to test for it
Before and after training, run the same handful of prompts from outside the trained task and compare the answers, not just the trained task's own score.

The hosted platform stops taking new jobs

How to notice it
A workflow built around retraining periodically can no longer submit a new job, though models already trained keep serving inference, because the maker wound the training service down without retiring what it produced.
How to test for it
Read the maker's own current guide for a notice like the one this page quotes before planning around a training service, not just the date the last job succeeded.

A reward the grader can game

How to notice it
A reinforcement-fine-tuned model's score against its own grader climbs while answers read by a person do not improve. This is the training-time version of the grader-hacking risk the site's evals topic covers for testing.
How to test for it
Hand-check a sample of the grader's own verdicts on the training data, the way a rubric grader's verdicts are hand-checked at eval time, rather than trusting the reward curve alone.

How to Evaluate It

This site’s own 60-question set does not score fine-tuning: the example on this page builds and checks a training file, and answers no question about the corpus, so the grading contract in evals/questions.json has nothing to check it against (see docs/EVALS.md). What a real fine-tuning job would still want measured is the same question set run twice (once on the base model, once on the fine-tuned one) so a claim of improvement is a before/after score on the same 60 questions rather than a training-time number alone, the way the evals page’s own rule requires. Alongside that: the leak check this page shows, run against the real split before spending anything on a job, and a handful of off-task prompts checked before and after, to catch the forgetting failure mode above.

Run it

What to monitor

Whether the fine-tuned model's behavior on real traffic still matches the last validation run; a growing gap usually means real inputs have drifted from the training data's shape, not that the model got worse on its own.

Cost at volume

Training is a one-time or periodic cost; the fine-tuned model can then cost less per call than a large general model prompted the long way, but only on the narrow task it was trained for. Traffic outside that task still needs the general model.

How it fails in production

The world moves and the training data doesn't. A model trained on last quarter's catalog answers confidently and wrong about this quarter's, with nothing in its own output flagging that its training predates the change.

What to log

The training data's version, the base model id, the method (supervised, DPO, reinforcement) and the date of the run, so a question about the model's behavior later can be traced to what it was actually trained on.

Try it

  1. Use it

    Find a product that advertises a small, fast, task-specific model. Look for whether its own page says LoRA/adapter, full fine-tuning, or doesn't say, and whether that silence changes how much you'd trust a claim that it 'matches' a bigger model.

  2. Build it

    Run python -m examples.adaptation --out .local/scratch/fine-tuning from the repo root, then open val.jsonl and count the lines. Change --val-fraction to 0.1 and run it again; does the count change the way split's docstring says it should?

  3. Either lane

    Pick two questions from evals/questions.json about the same appliance and decide whether they're close enough that training on one and validating on the other would leak. Run the leak check and see whether your judgment matches leaked_questions's.

How it connects

Before, after and instead of this

Pages that need this one

Optional: products, tools, and models

7 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 1 more examples
  • Unsloth Tool or framework · Unsloth

    Fine-tuning library

    Checked 09/18/2026
In practice

Learn a recurring response format

Train on reviewed input-output pairs and check the adapted model on examples excluded from training.

Out there

Named products, tools and models

Tools8
  • Axolotlopen source · fine-tuning library
  • MLXApple · training and inference on Apple hardware
  • OpenAI fine-tuningOpenAI · hosted fine-tuningRetired 2026
  • PEFTHugging Face · LoRA and other adapters
  • Together AI fine-tuningTogether AI · hosted fine-tuning
  • TransformersHugging Face · model library
  • TRLHugging Face · fine-tuning library
  • UnslothUnsloth · fine-tuning library

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. LoRA: Low-Rank Adaptation of Large Language Models · arXiv (Microsoft), 06/17/2021 (accessed 09/19/2026)
  2. Fine-tuning: overview · Together AI (accessed 09/19/2026)
  3. Fine-tuning: data preparation · Together AI (accessed 09/19/2026)
  4. Model optimization · OpenAI (API documentation) (accessed 09/19/2026)
  5. Supervised fine-tuning · OpenAI (API documentation) (accessed 09/19/2026)
  6. Reinforcement fine-tuning · OpenAI (API documentation) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page