examples/prompt_optimization is a minimal version of what a DSPy optimizer does, scoped down to
one thing: search over a fixed list of whole system prompts, using
distillation’s own grade_exact as the metric,
against the site’s 32 exact-graded questions. What DSPy’s own optimizers actually search over is
usually finer-grained: its documentation lists “synthesizing good few-shot examples for every
module,” “proposing and intelligently exploring better natural-language instructions for every
prompt,” and “building datasets for your modules and using them to finetune the LM weights” as
three different things an optimizer can tune[2]. This example only ever swaps the whole
instruction, never touches an example or a weight.
The split is the part worth reading closely:
examples/prompt_optimization/run.py · lines 103–177
def run(
tracer: Tracer,
model: Model,
*,
questions_path: Path = DEFAULT_QUESTIONS_PATH,
instructions: list[str] | None = None,
held_out_fraction: float = 0.25,
max_questions: int | None = None,
seed: int = 0,
) -> OptimizationResult:
instructions = instructions if instructions is not None else CANDIDATE_INSTRUCTIONS
questions = load_exact_questions(questions_path)
loaded = len(questions)
if max_questions is not None:
if max_questions < 1:
raise ValueError(f"max_questions={max_questions} leaves no questions to search over")
questions = questions[:max_questions]
detail = f"{len(questions)} questions"
if len(questions) < loaded:
detail = f"{len(questions)} of {loaded} questions, bounded by max_questions={max_questions}"
tracer.record(kind="code", decided_by="code", title="Load exact-graded questions", detail=detail)
dev, held_out = split_dev_held_out(questions, held_out_fraction=held_out_fraction, seed=seed)
if not dev:
# Selecting on an empty development split is not selection: every candidate ties at zero
# and `max` returns the first one, which would then be reported with a held-out score as
# though a search had chosen it. Fail here instead of returning a meaningless winner.
raise ValueError(
f"held_out_fraction={held_out_fraction} leaves no development questions to select on"
)
tracer.record(
kind="code",
decided_by="code",
title="Split into a development set and a held-out set",
detail=f"{len(dev)} development, {len(held_out)} held-out",
)
candidates: list[CandidateScore] = []
for instruction in instructions:
correct, total = _score(model, instruction, dev, tracer, phase="select")
candidates.append(CandidateScore(instruction=instruction, dev_correct=correct, dev_total=total))
tracer.record(
kind="code",
decided_by="code",
title="Score one candidate on the development split",
detail=f"{correct}/{total}",
)
# `max` keeps the first of equal scores, so a tie resolves to the earliest candidate in the
# list. That is a deterministic rule rather than a judgment: a run whose candidates all tie
# has selected nothing, and its "winner" is list order.
best = max(candidates, key=lambda c: c.dev_score)
tied = [c.instruction for c in candidates if c.dev_correct == best.dev_correct]
tracer.record(
kind="code",
decided_by="code",
title="Select the candidate with the best development score",
detail=f"{best.dev_score:.2f} on the development split"
+ (f"; {len(tied)} candidates tied, first in list order kept" if len(tied) > 1 else ""),
)
held_correct, held_total = _score(model, best.instruction, held_out, tracer, phase="report")
tracer.record(
kind="code",
decided_by="code",
title="Score the selected candidate on the held-out split",
detail=f"{held_correct}/{held_total}",
)
return OptimizationResult(
candidates=candidates,
selected=best.instruction,
held_out_correct=held_correct,
held_out_total=held_total,
)
Every candidate is scored on the development split; the highest score is selected; only then is
that one candidate scored on the held-out split. A development number answers “which candidate
looked best while we were choosing.” The held-out number answers “how good is the one we picked,”
and those are different questions.
The separation is worth attacking rather than believing, because a leak would change no number a
reader could see. Four properties hold it up, and each is a test. The split is a deterministic
partition for a given seed: the same seed gives the same two lists, every question lands in exactly
one of them, and the held-out side is never empty even at a fraction that rounds to zero. Every
held-out question is asked exactly once, after selection has finished and only under the winning
instruction: the test asserts the ordering, not just the instruction, since a held-out question
scored early and re-scored later would still have leaked. Every candidate is scored on the same
development questions. And ties resolve to the first candidate in list order, which means a run
where everything ties has selected nothing: python -m examples.prompt_optimization --model stub
does exactly that, scoring every candidate 0/24 against the echoing stub, so the “winner” is
list position. --model stub:scripted --max-questions 6 is the other case, bounded to a run small
enough to read: three candidates scored on the same four development questions, two of them taking
0/4 and one taking 4/4, which then confirms 2/2 on the two questions held back. A
held_out_fraction that would leave no development questions raises instead of
returning, because selecting on nothing and then printing a held-out score reads precisely like a
search that worked.
Size is the honest limitation. DSPy’s own guidance for a longer optimization run with its
MIPROv2 optimizer is to use it when you “have enough data (e.g. 200 examples or more to prevent
overfitting)”[2]: scoped to that one optimizer’s longer search mode, not a rule for
every method. This example splits 32 questions 24/8, enough to demonstrate the discipline and far
short of enough to trust a winner.