Nothing on this site trains a model: no example here calls a fine-tuning API, and no result file
exists for any adaptation technique. What a builder can do without one is prepare the data a
fine-tuning job would actually need, and check that it is not broken before spending anything on
a training run.
The example turns the site’s own 60-question set into a small supervised fine-tuning file. Each
question becomes one line in chat format (a system/user/assistant message list), the shape
Together AI’s fine-tuning data preparation guide documents: each message has “a role (system,
user, or assistant) and content”, and a conversation “must start with system or user
and alternate user and assistant afterwards”[2]. The same guide recommends holding
out a validation file rather than training on everything, so training progress can be checked
against examples the run never saw[2]; the example splits the 60 questions into a
training file and a validation file with a fixed random seed, so the same seed always produces
the same split.
examples/adaptation/run.py · lines 61–66
def load_examples(questions_path: Path = DEFAULT_QUESTIONS_PATH) -> list[Example]:
"""Every question in the set, as a training example. Every question carries a plain-language
`answer` even when its kind is `unanswerable` (the correct completion there is the model
saying so), so nothing is filtered out by kind."""
data = json.loads(questions_path.read_text(encoding="utf-8"))
return [Example(id=q["id"], question=q["question"], answer=q["answer"]) for q in data["questions"]]
A held-out split only means something if training never saw the held-out questions under a
different guise. leaked_questions checks every validation question’s normalized text against
the training file’s, and reports any that show up in both:
examples/adaptation/run.py · lines 83–92
def leaked_questions(train: list[Example], val: list[Example]) -> list[str]:
"""Val-set ids whose normalized question text also appears in the train set. A held-out split
is only worth anything if training never saw the same question under a different id.
This catches only questions that are identical once normalized. A paraphrase in genuinely
different words ("how long is the warranty" against "what is the warranty period") is a
near-duplicate this equality test cannot see; catching those needs a similarity measure, and
an embedding of each question is the usual one."""
train_texts = {_normalize(ex.question) for ex in train}
return sorted(ex.id for ex in val if _normalize(ex.question) in train_texts)
_normalize lowercases, drops punctuation and collapses whitespace, so two questions that differ
only in styling count as one. On this set the check comes back empty: all 60 questions normalize
to distinct strings. The test file proves it catches a leak by putting the same question,
restyled, on opposite sides of the split, and proves what it misses, by letting a genuine
paraphrase through. Equality over normalized text is the floor. Catching paraphrases needs a
similarity measure instead, usually an embedding of each question, and that is the gap that
starts to matter on a larger set built partly from generated questions: the job Argilla’s
distilabel describes itself as built for, “a framework for synthetic data and AI
feedback”[6].
Two techniques this page covers have no data-preparation step to show. Reinforcement fine-tuning
trains against a grader’s score rather than fixed example answers: OpenAI’s own guide says the
method “samples several responses per prompt, scores them with the grader, and applies
policy-gradient updates based on those rewards”[4], so there is no training file to
build at all. Automated prompt optimization tunes a prompt or its few-shot examples against a
metric instead of retraining weights; DSPy’s own description is “algorithms for optimizing their
prompts and weights” so a program does not depend on “brittle prompts”[5], which is
closer to what this site’s own eval loop would drive than to a training file.
Which of these a maker currently sells changes faster than the methods do. Both OpenAI guides
cited here carried this notice when they were read for this page: “OpenAI is winding down the
fine-tuning platform. The platform is no longer accessible to new users, but existing users of
the fine-tuning platform will be able to create training jobs for the coming
months”[3][4].
Fine-tuned models stay available for inference until their base models are deprecated, the same
notice adds; the reinforcement fine-tuning guide also limits that method
to one reasoning model id[4]. Together AI’s fine-tuning documentation carries no such
notice[2]. Read the maker’s own page before planning around a service; the methods
outlast the platforms that sell them.