Nothing on this site trains a model: no example here calls a fine-tuning API, on this page or on
the adaptation page it hangs from. What
examples/adaptation builds instead is the file a supervised fine-tuning job would actually need:
a chat-format JSONL split into training and validation questions from the site’s own 60-question
set, with a check that no question leaks across the split. The adaptation page shows the part that
turns a question into a training example and the leak check itself; this page shows the split those
two functions sit between.
split shuffles with a fixed seed and cuts a validation fraction off the top, so the same seed
always produces the same partition:
examples/adaptation/run.py · lines 95–99
def split(examples: list[Example], *, val_fraction: float, seed: int) -> tuple[list[Example], list[Example]]:
shuffled = list(examples)
random.Random(seed).shuffle(shuffled)
val_count = max(1, round(len(shuffled) * val_fraction))
return shuffled[val_count:], shuffled[:val_count] # train, val
The default, 20%, exists for the same reason Together AI’s own fine-tuning data preparation guide
gives. Its instructions for carving “a validation set out of a single JSONL file” continue:
“Then pass both files to the job and set n_evals above 0:” and, further on, “The model evaluates
against the validation set at the specified intervals” during training[3]. A held-out set
is only useful if something is scored against it while the job runs, not just kept aside.
run is the whole pipeline in order: load the questions, split them, write both files, then check
the split for a leak.
examples/adaptation/run.py · lines 108–139
def run(
tracer: Tracer,
*,
questions_path: Path = DEFAULT_QUESTIONS_PATH,
out_dir: Path,
val_fraction: float = 0.2,
seed: int = 0,
) -> BuildResult:
examples = load_examples(questions_path)
tracer.record(kind="code", decided_by="code", title="Load questions as chat examples", detail=f"{len(examples)} examples")
train, val = split(examples, val_fraction=val_fraction, seed=seed)
tracer.record(
kind="code",
decided_by="code",
title="Shuffle and split into train and validation",
detail=f"{len(train)} train, {len(val)} val, seed={seed}",
)
train_path, val_path = out_dir / "train.jsonl", out_dir / "val.jsonl"
_write_jsonl(train_path, train)
_write_jsonl(val_path, val)
tracer.record(kind="code", decided_by="code", title="Write JSONL files", detail=f"{train_path.name}, {val_path.name}")
leaked = leaked_questions(train, val)
tracer.record(
kind="code",
decided_by="code",
title="Check for leaked questions between splits",
detail=", ".join(leaked) or "none",
)
return BuildResult(train=train, val=val, train_path=train_path, val_path=val_path, leaked=leaked)
Every step it records is decided_by: "code": the split is a fixed shuffle-and-cut, not a choice
a model makes, so this example contributes zero model-decided steps, the same as any other data
preparation step.
Preference tuning and reinforcement fine-tuning have no file to build the way supervised
fine-tuning does, because neither trains against one fixed right answer. Together AI’s own
overview separates two of its training methods on exactly that line: supervised fine-tuning trains
“on demonstration data with one target completion per example,” while preference fine-tuning is
described as, “Align a model with rankings over preferred and dispreferred responses using
DPO.”[2] OpenAI documents its own DPO option in the same shape, telling a reader to
“Provide both a correct and incorrect example response for a prompt. Indicate the correct response
to help the model perform better.”[4] It lists three model ids the method is available
for today: gpt-4.1-2025-04-14, gpt-4.1-mini-2025-04-14 and gpt-4.1-nano-2025-04-14.
Reinforcement fine-tuning goes further still: OpenAI’s own guide says that during training the
platform “samples several responses per prompt, scores them with the grader, and applies
policy-gradient updates based on those rewards.”[6] The method “is supported on o-series
reasoning models only, and currently only for o4-mini”[6]. Both are OpenAI’s terms for
OpenAI’s platform, read September 19, 2026; another maker’s DPO offering is its own to describe. A
training file for either would be pairs or grader code, not the chat JSONL this example writes.
Self-run alternatives exist for anyone who would rather not depend on a hosted platform at all:
Hugging Face’s Transformers, TRL and PEFT, Unsloth, Axolotl and Apple’s MLX all train adapters or
full weights on your own hardware, at the cost of running the training yourself.