# Prompt optimization

_Topics at every level · sourced_

Letting a program search for better prompts against a test set.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a search over prompt candidates and compare their performance beyond the examples used to select them. Inspect the difference between improving a score and improving the actual task.

**Assumptions:** The objective and development set shape what gets optimized. A narrow score may reward behavior that is unhelpful elsewhere.

**Design choices:** Use a limited candidate search, a meaningful baseline, and untouched evaluation cases. Include cost or complexity when they affect deployment value.

**Request:** Search extraction prompts without overfitting the final test set.

**Starting evidence:** Development D and sealed test T. P1 fills all fields; P2 preserves unknowns.

**Action and control:** Compare prompts on D using factual criteria; select before opening T.

**Stage records (authored, not executed):**

### Input record

Development D and sealed test T. P1 fills all fields; P2 preserves unknowns.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use a limited candidate search, a meaningful baseline, and untouched evaluation cases. Include cost or complexity when they affect deployment value.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Compare prompts on D using factual criteria; select before opening T.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Record the selected prompt and development evidence; final testing stays separate. No measured gain invented.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Candidate prompts, development scores, a grader loophole, and final held-out comparison with versioned prompts.

If the result falls short:
If a winning prompt fails on new cases, inspect overfitting and hidden assumptions. Do not keep modifying the final test set to preserve the apparent win.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for extraction, classification, or other repeated prompts. A manually improved prompt may be sufficient when the task or dataset is still changing.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Record the selected prompt and development evidence; final testing stays separate. No measured gain invented.

**Change something — Grader rewards every nonempty field:** Invented values win the flawed objective. Repair the grader and re-evaluate.

**Decision:** Should a prompt be accepted just because its score rose?

**Answer:** No; inspect the objective and held-out behavior.

**Why:** Optimization can exploit the grader or overfit development examples; preserve an untouched test set.

**Review criteria:** Candidate prompts, development scores, a grader loophole, and final held-out comparison with versioned prompts.

**Recovery:** If a winning prompt fails on new cases, inspect overfitting and hidden assumptions. Do not keep modifying the final test set to preserve the apparent win.

**Adapt it:** Use this for extraction, classification, or other repeated prompts. A manually improved prompt may be sufficient when the task or dataset is still changing.

A program searches for a better prompt against a measured score, instead of a person hand-editing
the wording. That is prompt optimization, also called automated prompt optimization, and it is the
one technique under [changing the model](/gradient_ascent/techniques/adaptation/) that changes no
weights at all: where [fine-tuning](/gradient_ascent/techniques/fine-tuning/) trains the
instruction in, this searches for a better one to send.

DSPy is the program this page quotes throughout. Its paper describes designing "a compiler that
will optimize any DSPy pipeline to maximize a given metric"[1], and its documentation
names three things an optimizer takes: the program itself, a metric, and a handful of training
inputs, which it says "may be very small (i.e., only 5 or 10 examples) and
incomplete"[2]. Two of those three have to exist before there is anything to
optimize, which is why this page assumes [evals](/gradient_ascent/techniques/evals/). An eval set
is not a nice-to-have here. It is the thing being optimized against, and every weakness in it is
inherited by whatever comes out.

This page is sourced, not measured: the search methods below come from their authors' own papers
and libraries, and no optimization run has happened here.

## Practical guidance

Chat apps increasingly ship an "improve this prompt" button. Anthropic's own version rewrites a
prompt in one pass using fixed techniques, then lets you keep adjusting it: "you can provide
feedback for Claude about what is and isn't working to further improve the prompt"[3].
That button compares no candidates and consults no score. It hands you a better first draft, not a
measurement of whether the new prompt actually works better than the old one.

To find that out, run the by-hand version yourself. Collect ten real questions you already know
the right answer to. Run all ten through your current prompt and mark which came back right. Run
the same ten through the rewritten prompt, unchanged otherwise, and mark which came back right.
Count both. A rewrite that gets seven right against the old prompt's eight is not an improvement,
however much better it reads on the page.

Some tools go a step further and let a person grade the outputs: Anthropic's own prompt evaluator
lets you "test your prompts under various scenarios"[3] and adds an "ideal output" column
so you can "grade model outputs on a 5-point scale"[3]. That is still a person grading
one prompt at a time, not a search comparing prompts by number; read that grade the same way you'd
read your own ten-question count, as one data point, not a verdict.

An optimized or "tuned" prompt someone hands you is also not permanent: it was fitted to one set of
test questions on one model, and switching models, including a new version from the same maker,
can make an old winner perform worse than the plain prompt it beat. If a prompt you rely on was
last checked against a model that has since changed, re-run your own ten questions before trusting
it again.

None of this is worth doing for a question you will ask once. Write it, read the answer, move on.
Set up the ten-question comparison only once you are about to keep a rewritten prompt running for a
while, since that is the point where "seems better" starts costing something if it turns out wrong.

## Implementation details

`examples/prompt_optimization` is a minimal version of what a DSPy optimizer does, scoped down to
one thing: search over a fixed list of whole system prompts, using
[distillation's own](/gradient_ascent/techniques/distillation/) `grade_exact` as the metric,
against the site's 32 exact-graded questions. What DSPy's own optimizers actually search over is
usually finer-grained: its documentation lists "synthesizing good few-shot examples for every
module," "proposing and intelligently exploring better natural-language instructions for every
prompt," and "building datasets for your modules and using them to finetune the LM weights" as
three different things an optimizer can tune[2]. This example only ever swaps the whole
instruction, never touches an example or a weight.

The split is the part worth reading closely:

`examples/prompt_optimization/run.py` (lines 103-177)

```python
def run(
    tracer: Tracer,
    model: Model,
    *,
    questions_path: Path = DEFAULT_QUESTIONS_PATH,
    instructions: list[str] | None = None,
    held_out_fraction: float = 0.25,
    max_questions: int | None = None,
    seed: int = 0,
) -> OptimizationResult:
    instructions = instructions if instructions is not None else CANDIDATE_INSTRUCTIONS
    questions = load_exact_questions(questions_path)
    loaded = len(questions)
    if max_questions is not None:
        if max_questions < 1:
            raise ValueError(f"max_questions={max_questions} leaves no questions to search over")
        questions = questions[:max_questions]
    detail = f"{len(questions)} questions"
    if len(questions) < loaded:
        detail = f"{len(questions)} of {loaded} questions, bounded by max_questions={max_questions}"
    tracer.record(kind="code", decided_by="code", title="Load exact-graded questions", detail=detail)

    dev, held_out = split_dev_held_out(questions, held_out_fraction=held_out_fraction, seed=seed)
    if not dev:
        # Selecting on an empty development split is not selection: every candidate ties at zero
        # and `max` returns the first one, which would then be reported with a held-out score as
        # though a search had chosen it. Fail here instead of returning a meaningless winner.
        raise ValueError(
            f"held_out_fraction={held_out_fraction} leaves no development questions to select on"
        )
    tracer.record(
        kind="code",
        decided_by="code",
        title="Split into a development set and a held-out set",
        detail=f"{len(dev)} development, {len(held_out)} held-out",
    )

    candidates: list[CandidateScore] = []
    for instruction in instructions:
        correct, total = _score(model, instruction, dev, tracer, phase="select")
        candidates.append(CandidateScore(instruction=instruction, dev_correct=correct, dev_total=total))
        tracer.record(
            kind="code",
            decided_by="code",
            title="Score one candidate on the development split",
            detail=f"{correct}/{total}",
        )

    # `max` keeps the first of equal scores, so a tie resolves to the earliest candidate in the
    # list. That is a deterministic rule rather than a judgment: a run whose candidates all tie
    # has selected nothing, and its "winner" is list order.
    best = max(candidates, key=lambda c: c.dev_score)
    tied = [c.instruction for c in candidates if c.dev_correct == best.dev_correct]
    tracer.record(
        kind="code",
        decided_by="code",
        title="Select the candidate with the best development score",
        detail=f"{best.dev_score:.2f} on the development split"
        + (f"; {len(tied)} candidates tied, first in list order kept" if len(tied) > 1 else ""),
    )

    held_correct, held_total = _score(model, best.instruction, held_out, tracer, phase="report")
    tracer.record(
        kind="code",
        decided_by="code",
        title="Score the selected candidate on the held-out split",
        detail=f"{held_correct}/{held_total}",
    )

    return OptimizationResult(
        candidates=candidates,
        selected=best.instruction,
        held_out_correct=held_correct,
        held_out_total=held_total,
    )
```

Every candidate is scored on the development split; the highest score is selected; only then is
that one candidate scored on the held-out split. A development number answers "which candidate
looked best while we were choosing." The held-out number answers "how good is the one we picked,"
and those are different questions.

The separation is worth attacking rather than believing, because a leak would change no number a
reader could see. Four properties hold it up, and each is a test. The split is a deterministic
partition for a given seed: the same seed gives the same two lists, every question lands in exactly
one of them, and the held-out side is never empty even at a fraction that rounds to zero. Every
held-out question is asked exactly once, after selection has finished and only under the winning
instruction: the test asserts the ordering, not just the instruction, since a held-out question
scored early and re-scored later would still have leaked. Every candidate is scored on the same
development questions. And ties resolve to the first candidate in list order, which means a run
where everything ties has selected nothing: `python -m examples.prompt_optimization --model stub`
does exactly that, scoring every candidate 0/24 against the echoing stub, so the "winner" is
list position. `--model stub:scripted --max-questions 6` is the other case, bounded to a run small
enough to read: three candidates scored on the same four development questions, two of them taking
0/4 and one taking 4/4, which then confirms 2/2 on the two questions held back. A
`held_out_fraction` that would leave no development questions raises instead of
returning, because selecting on nothing and then printing a held-out score reads precisely like a
search that worked.

Size is the honest limitation. DSPy's own guidance for a longer optimization run with its
`MIPROv2` optimizer is to use it when you "have enough data (e.g. 200 examples or more to prevent
overfitting)"[2]: scoped to that one optimizer's longer search mode, not a rule for
every method. This example splits 32 questions 24/8, enough to demonstrate the discipline and far
short of enough to trust a winner.

## When you do not need this

There is a prerequisite here that rules most cases out on its own: no eval set, no search. A
program that picks a prompt by score cannot run without a metric and examples to run it against,
so if you do not already have an eval set you trust, the work in front of you is
[building one](/gradient_ascent/techniques/evals/), and that work usually improves the prompt by
itself: writing down what a good answer looks like is most of saying what you want. Start where
[prompt engineering](/gradient_ascent/techniques/prompt-engineering/) starts: one prompt written
by hand, checked against a handful of real cases.

Search becomes worth its cost at the point where comparing candidates by hand is the slow part:
the same prompt sent often, an eval set someone has already sampled and trusts, and more variants
worth trying than a person will sit through. Short of that, a machine comparing dozens of prompts
against a set nobody has checked will confidently hand you the one that best fits its flaws.

## Failure modes

### The winning candidate never faces held-out data

- **How to notice it:** A reported score is the same number the search used to choose the candidate in the first place, so it measures how well the search fit that one set, not how the candidate performs elsewhere.
- **How to test for it:** Check whether the reported number came from the same examples the search compared candidates on. If so, it is a training-time number, not a held-out one, whatever it is called on the page.

### Too little data for the amount of search

- **How to notice it:** A search tries many candidates against a small example set, and the winner's score is really noise from that small set rather than a real difference between candidates.
- **How to test for it:** Compare the number of candidates tried against the number of examples scored on; DSPy's own guidance scopes its 200-example recommendation to one optimizer's longer search mode, which is a useful reference point even for a different search.

### Metric mismatch between what is optimized and what is wanted

- **How to notice it:** The search maximizes exactly the metric it was given, and the metric turns out to reward something narrower than what the prompt was actually supposed to do well.
- **How to test for it:** Read a sample of the highest-scoring candidate's actual outputs, not just its score, and check whether a person would call them good for the real task.

### A one-shot rewrite mistaken for a search

- **How to notice it:** A prompt-improvement tool rewrites a prompt once using fixed techniques, and the result is passed on as though something had compared it against alternatives and measured the difference.
- **How to test for it:** Ask what scored it, and on what. A rewrite is a draft: it needs the same check by hand that any prompt you wrote yourself would need, and a grade a person gave one prompt is not a comparison between two.

### The model changed and the prompt did not

- **How to notice it:** An optimized prompt keeps running after the model behind it is upgraded or swapped, still carrying a held-out score that was measured on the old one. Instructions tuned around one model's habits can be neutral or harmful on the next.
- **How to test for it:** Re-score the current prompt on the held-out split against the new model before the switch, and re-run the search if the number moved. Record the model id beside every score so this question can be asked at all.

### The eval set the search runs against is the problem

- **How to notice it:** The search finds a real, generalizable improvement against a flawed or unrepresentative eval set, and the improvement does not show up once the prompt meets real traffic.
- **How to test for it:** Before trusting an optimization result, apply the same checks the evals page describes to the set itself: is it representative of real questions, and has anyone checked a sample by hand?

## How to Evaluate It

Running the site's eval runner over this example would be circular: the example already uses 32 of
those 60 questions as the thing it searches and reports against (see `docs/EVALS.md`). Its own
held-out score is the measurement, and it means something only because of where the number came
from: data the choice never touched, which is the rule [evals](/gradient_ascent/techniques/evals/)
sets for every before/after on this site.

Optimizing a real prompt against the full 60 works the same way: cut a slice off first, search on
the rest, score the winner on the slice, and report the two numbers separately with the model id
beside them. One number labeled "after optimization," with no statement of which questions
produced it, is not a result.

## Run it

**What to monitor.** The gap between the development score and the held-out score over repeated
  optimization runs; a gap that keeps growing usually means the eval set is too small, too
  similar to what the search has already seen, or both.

**Cost at volume.** Cost scales with candidates tried times examples scored per candidate, on the
  development split; the held-out check adds one more full pass, once, for whichever candidate
  won. A wider search is a multiplier on the development side only.

**How it fails in production.** Two drifts, and the prompt notices neither. Real traffic moves away from the
  shape of the questions the search ran on, and the model behind the prompt is upgraded to one
  that reads the same instruction differently. In both cases a held-out score measured months ago
  keeps being quoted as though it still described the system.

**What to log.** Every candidate tried and its development score, which one was selected and why, the
  held-out score, and the eval set's own version, so a later question about why this prompt was
  chosen can be answered by rereading a record instead of rerunning the search.

## Try it

1. **Use it.** Find a claim that a prompt was 'optimized' or 'tuned' for a task, and try to answer two questions from the page alone: how many examples was it scored on, and were any of them kept back from the process that picked it? Most pages answer neither, and the second one is the one that decides what the number means.
2. **Build it.** Run python -m examples.prompt_optimization --model stub from the repo root and read the three development scores. Then open examples/prompt_optimization/run.py and add a fourth instruction at the TOP of CANDIDATE_INSTRUCTIONS. Run it again: every score is still 0/24, the held-out score is still 0/8, and the selected candidate has changed. Selection by list position is what a tie actually is.
3. **Either lane.** Write two short system prompts for a task you know well, and by hand, run each against five questions you already know the right answer to. Keep two of those five aside before you look at either prompt's answers, then score only the remaining three to pick a winner, and check the winner's score on the two you kept aside. Did the held-out pair agree with your pick?


## Sources

1. [DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines](https://arxiv.org/abs/2310.03714) — arXiv (Khattab et al., Stanford), 2023-10-05 (accessed 2026-09-19)
2. [DSPy Optimizers (formerly Teleprompters)](https://github.com/stanfordnlp/dspy/blob/main/docs/docs/learn/optimization/optimizers.md) — DSPy (Stanford NLP), documentation source (accessed 2026-09-19)
3. [Improve your prompts in the developer console](https://claude.com/blog/prompt-improver) — Anthropic, 2024-10-14 (accessed 2026-09-19)


Last reviewed 2026-09-19.
