Topics at every level

Prompt optimization

Letting a program search for better prompts against a test set.

Sourced

Concept at a glance

Search for a better prompt against a defined test.

Feedback loopConceptual illustration
Search for a better prompt against a defined test.Task + dev cases leads to Candidate prompt. Candidate prompt leads to Evaluate. Evaluate leads to Select + revise. Select + revise leads to Candidate prompt as feedback. Keep a held-out set separate from the examples used to choose the prompt.Task + dev casesDefine the scoring targetCandidate promptTry a new instructionEvaluateScore on development casesSelect + reviseUse results for the nexttrialSearch for a better prompt against a defined test.Task + dev cases leads to Candidate prompt. Candidate prompt leads to Evaluate. Evaluate leads to Select + revise. Select + revise leads to Candidate prompt as feedback. Keep a held-out set separate from the examples used to choose the prompt.Task + dev casesDefine the scoring targetCandidate promptTry a new instructionEvaluateScore on development casesSelect + reviseUse results for the nexttrial

Ending or continuingSelect using development cases, then measure once on held-out cases.

Read the connections in words
  • Task + dev cases → Candidate prompt: Try a new instruction.
  • Candidate prompt → Evaluate: Score on development cases.
  • Evaluate → Select + revise: Use results for the next trial.
  • Select + revise → Candidate prompt: feedback informs another turn.
Key idea

Keep a held-out set separate from the examples used to choose the prompt.

A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Prompt optimization: see it in practice.

Searching prompt variants against an objective using a development set.

What you’ll walk through

Follow a search over prompt candidates and compare their performance beyond the examples used to select them. Inspect the difference between improving a score and improving the actual task.

The task in this version

Search extraction prompts without overfitting the final test set.

What you’ll learn to check

Candidate prompts, development scores, a grader loophole, and final held-out comparison with versioned prompts.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Search extraction prompts without overfitting the final test set.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Development D and sealed test T. P1 fills all fields; P2 preserves unknowns.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

The objective and development set shape what gets optimized. A narrow score may reward behavior that is unhelpful elsewhere.

1 / 6

Apply this to your project

Describe your task to your own model and use Prompt optimization as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

A program searches for a better prompt against a measured score, instead of a person hand-editing the wording. That is prompt optimization, also called automated prompt optimization, and it is the one technique under changing the model that changes no weights at all: where fine-tuning trains the instruction in, this searches for a better one to send.

DSPy is the program this page quotes throughout. Its paper describes designing “a compiler that will optimize any DSPy pipeline to maximize a given metric”[1], and its documentation names three things an optimizer takes: the program itself, a metric, and a handful of training inputs, which it says “may be very small (i.e., only 5 or 10 examples) and incomplete”[2]. Two of those three have to exist before there is anything to optimize, which is why this page assumes evals. An eval set is not a nice-to-have here. It is the thing being optimized against, and every weakness in it is inherited by whatever comes out.

This page is sourced, not measured: the search methods below come from their authors’ own papers and libraries, and no optimization run has happened here.

Practical guidance

Chat apps increasingly ship an “improve this prompt” button. Anthropic’s own version rewrites a prompt in one pass using fixed techniques, then lets you keep adjusting it: “you can provide feedback for Claude about what is and isn’t working to further improve the prompt”[3]. That button compares no candidates and consults no score. It hands you a better first draft, not a measurement of whether the new prompt actually works better than the old one.

To find that out, run the by-hand version yourself. Collect ten real questions you already know the right answer to. Run all ten through your current prompt and mark which came back right. Run the same ten through the rewritten prompt, unchanged otherwise, and mark which came back right. Count both. A rewrite that gets seven right against the old prompt’s eight is not an improvement, however much better it reads on the page.

Some tools go a step further and let a person grade the outputs: Anthropic’s own prompt evaluator lets you “test your prompts under various scenarios”[3] and adds an “ideal output” column so you can “grade model outputs on a 5-point scale”[3]. That is still a person grading one prompt at a time, not a search comparing prompts by number; read that grade the same way you’d read your own ten-question count, as one data point, not a verdict.

An optimized or “tuned” prompt someone hands you is also not permanent: it was fitted to one set of test questions on one model, and switching models, including a new version from the same maker, can make an old winner perform worse than the plain prompt it beat. If a prompt you rely on was last checked against a model that has since changed, re-run your own ten questions before trusting it again.

None of this is worth doing for a question you will ask once. Write it, read the answer, move on. Set up the ten-question comparison only once you are about to keep a rewritten prompt running for a while, since that is the point where “seems better” starts costing something if it turns out wrong.

Implementation details

examples/prompt_optimization is a minimal version of what a DSPy optimizer does, scoped down to one thing: search over a fixed list of whole system prompts, using distillation’s own grade_exact as the metric, against the site’s 32 exact-graded questions. What DSPy’s own optimizers actually search over is usually finer-grained: its documentation lists “synthesizing good few-shot examples for every module,” “proposing and intelligently exploring better natural-language instructions for every prompt,” and “building datasets for your modules and using them to finetune the LM weights” as three different things an optimizer can tune[2]. This example only ever swaps the whole instruction, never touches an example or a weight.

The split is the part worth reading closely:

examples/prompt_optimization/run.py · lines 103–177
def run(
    tracer: Tracer,
    model: Model,
    *,
    questions_path: Path = DEFAULT_QUESTIONS_PATH,
    instructions: list[str] | None = None,
    held_out_fraction: float = 0.25,
    max_questions: int | None = None,
    seed: int = 0,
) -> OptimizationResult:
    instructions = instructions if instructions is not None else CANDIDATE_INSTRUCTIONS
    questions = load_exact_questions(questions_path)
    loaded = len(questions)
    if max_questions is not None:
        if max_questions < 1:
            raise ValueError(f"max_questions={max_questions} leaves no questions to search over")
        questions = questions[:max_questions]
    detail = f"{len(questions)} questions"
    if len(questions) < loaded:
        detail = f"{len(questions)} of {loaded} questions, bounded by max_questions={max_questions}"
    tracer.record(kind="code", decided_by="code", title="Load exact-graded questions", detail=detail)

    dev, held_out = split_dev_held_out(questions, held_out_fraction=held_out_fraction, seed=seed)
    if not dev:
        # Selecting on an empty development split is not selection: every candidate ties at zero
        # and `max` returns the first one, which would then be reported with a held-out score as
        # though a search had chosen it. Fail here instead of returning a meaningless winner.
        raise ValueError(
            f"held_out_fraction={held_out_fraction} leaves no development questions to select on"
        )
    tracer.record(
        kind="code",
        decided_by="code",
        title="Split into a development set and a held-out set",
        detail=f"{len(dev)} development, {len(held_out)} held-out",
    )

    candidates: list[CandidateScore] = []
    for instruction in instructions:
        correct, total = _score(model, instruction, dev, tracer, phase="select")
        candidates.append(CandidateScore(instruction=instruction, dev_correct=correct, dev_total=total))
        tracer.record(
            kind="code",
            decided_by="code",
            title="Score one candidate on the development split",
            detail=f"{correct}/{total}",
        )

    # `max` keeps the first of equal scores, so a tie resolves to the earliest candidate in the
    # list. That is a deterministic rule rather than a judgment: a run whose candidates all tie
    # has selected nothing, and its "winner" is list order.
    best = max(candidates, key=lambda c: c.dev_score)
    tied = [c.instruction for c in candidates if c.dev_correct == best.dev_correct]
    tracer.record(
        kind="code",
        decided_by="code",
        title="Select the candidate with the best development score",
        detail=f"{best.dev_score:.2f} on the development split"
        + (f"; {len(tied)} candidates tied, first in list order kept" if len(tied) > 1 else ""),
    )

    held_correct, held_total = _score(model, best.instruction, held_out, tracer, phase="report")
    tracer.record(
        kind="code",
        decided_by="code",
        title="Score the selected candidate on the held-out split",
        detail=f"{held_correct}/{held_total}",
    )

    return OptimizationResult(
        candidates=candidates,
        selected=best.instruction,
        held_out_correct=held_correct,
        held_out_total=held_total,
    )

Every candidate is scored on the development split; the highest score is selected; only then is that one candidate scored on the held-out split. A development number answers “which candidate looked best while we were choosing.” The held-out number answers “how good is the one we picked,” and those are different questions.

The separation is worth attacking rather than believing, because a leak would change no number a reader could see. Four properties hold it up, and each is a test. The split is a deterministic partition for a given seed: the same seed gives the same two lists, every question lands in exactly one of them, and the held-out side is never empty even at a fraction that rounds to zero. Every held-out question is asked exactly once, after selection has finished and only under the winning instruction: the test asserts the ordering, not just the instruction, since a held-out question scored early and re-scored later would still have leaked. Every candidate is scored on the same development questions. And ties resolve to the first candidate in list order, which means a run where everything ties has selected nothing: python -m examples.prompt_optimization --model stub does exactly that, scoring every candidate 0/24 against the echoing stub, so the “winner” is list position. --model stub:scripted --max-questions 6 is the other case, bounded to a run small enough to read: three candidates scored on the same four development questions, two of them taking 0/4 and one taking 4/4, which then confirms 2/2 on the two questions held back. A held_out_fraction that would leave no development questions raises instead of returning, because selecting on nothing and then printing a held-out score reads precisely like a search that worked.

Size is the honest limitation. DSPy’s own guidance for a longer optimization run with its MIPROv2 optimizer is to use it when you “have enough data (e.g. 200 examples or more to prevent overfitting)”[2]: scoped to that one optimizer’s longer search mode, not a rule for every method. This example splits 32 questions 24/8, enough to demonstrate the discipline and far short of enough to trust a winner.

When you do not need this

There is a prerequisite here that rules most cases out on its own: no eval set, no search. A program that picks a prompt by score cannot run without a metric and examples to run it against, so if you do not already have an eval set you trust, the work in front of you is building one, and that work usually improves the prompt by itself: writing down what a good answer looks like is most of saying what you want. Start where prompt engineering starts: one prompt written by hand, checked against a handful of real cases.

Search becomes worth its cost at the point where comparing candidates by hand is the slow part: the same prompt sent often, an eval set someone has already sampled and trusts, and more variants worth trying than a person will sit through. Short of that, a machine comparing dozens of prompts against a set nobody has checked will confidently hand you the one that best fits its flaws.

Failure modes

The winning candidate never faces held-out data

How to notice it
A reported score is the same number the search used to choose the candidate in the first place, so it measures how well the search fit that one set, not how the candidate performs elsewhere.
How to test for it
Check whether the reported number came from the same examples the search compared candidates on. If so, it is a training-time number, not a held-out one, whatever it is called on the page.

Too little data for the amount of search

How to notice it
A search tries many candidates against a small example set, and the winner's score is really noise from that small set rather than a real difference between candidates.
How to test for it
Compare the number of candidates tried against the number of examples scored on; DSPy's own guidance scopes its 200-example recommendation to one optimizer's longer search mode, which is a useful reference point even for a different search.

Metric mismatch between what is optimized and what is wanted

How to notice it
The search maximizes exactly the metric it was given, and the metric turns out to reward something narrower than what the prompt was actually supposed to do well.
How to test for it
Read a sample of the highest-scoring candidate's actual outputs, not just its score, and check whether a person would call them good for the real task.

A one-shot rewrite mistaken for a search

How to notice it
A prompt-improvement tool rewrites a prompt once using fixed techniques, and the result is passed on as though something had compared it against alternatives and measured the difference.
How to test for it
Ask what scored it, and on what. A rewrite is a draft: it needs the same check by hand that any prompt you wrote yourself would need, and a grade a person gave one prompt is not a comparison between two.

The model changed and the prompt did not

How to notice it
An optimized prompt keeps running after the model behind it is upgraded or swapped, still carrying a held-out score that was measured on the old one. Instructions tuned around one model's habits can be neutral or harmful on the next.
How to test for it
Re-score the current prompt on the held-out split against the new model before the switch, and re-run the search if the number moved. Record the model id beside every score so this question can be asked at all.

The eval set the search runs against is the problem

How to notice it
The search finds a real, generalizable improvement against a flawed or unrepresentative eval set, and the improvement does not show up once the prompt meets real traffic.
How to test for it
Before trusting an optimization result, apply the same checks the evals page describes to the set itself: is it representative of real questions, and has anyone checked a sample by hand?

How to Evaluate It

Running the site’s eval runner over this example would be circular: the example already uses 32 of those 60 questions as the thing it searches and reports against (see docs/EVALS.md). Its own held-out score is the measurement, and it means something only because of where the number came from: data the choice never touched, which is the rule evals sets for every before/after on this site.

Optimizing a real prompt against the full 60 works the same way: cut a slice off first, search on the rest, score the winner on the slice, and report the two numbers separately with the model id beside them. One number labeled “after optimization,” with no statement of which questions produced it, is not a result.

Run it

What to monitor

The gap between the development score and the held-out score over repeated optimization runs; a gap that keeps growing usually means the eval set is too small, too similar to what the search has already seen, or both.

Cost at volume

Cost scales with candidates tried times examples scored per candidate, on the development split; the held-out check adds one more full pass, once, for whichever candidate won. A wider search is a multiplier on the development side only.

How it fails in production

Two drifts, and the prompt notices neither. Real traffic moves away from the shape of the questions the search ran on, and the model behind the prompt is upgraded to one that reads the same instruction differently. In both cases a held-out score measured months ago keeps being quoted as though it still described the system.

What to log

Every candidate tried and its development score, which one was selected and why, the held-out score, and the eval set's own version, so a later question about why this prompt was chosen can be answered by rereading a record instead of rerunning the search.

Try it

  1. Use it

    Find a claim that a prompt was 'optimized' or 'tuned' for a task, and try to answer two questions from the page alone: how many examples was it scored on, and were any of them kept back from the process that picked it? Most pages answer neither, and the second one is the one that decides what the number means.

  2. Build it

    Run python -m examples.prompt_optimization --model stub from the repo root and read the three development scores. Then open examples/prompt_optimization/run.py and add a fourth instruction at the TOP of CANDIDATE_INSTRUCTIONS. Run it again: every score is still 0/24, the held-out score is still 0/8, and the selected candidate has changed. Selection by list position is what a tie actually is.

  3. Either lane

    Write two short system prompts for a task you know well, and by hand, run each against five questions you already know the right answer to. Keep two of those five aside before you look at either prompt's answers, then score only the remaining three to pick a winner, and check the winner's score on the two you kept aside. Did the held-out pair agree with your pick?

How it connects

Before, after and instead of this

Read first

Often used with

Optional: products, tools, and models

1 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Search for clearer extraction instructions

Score candidate prompts on development examples, choose a winner, and evaluate it on untouched cases.

Out there

Named products, tools and models

Tools1
  • DSPyStanford NLP · prompt programs and optimizers

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines · arXiv (Khattab et al., Stanford), 10/05/2023 (accessed 09/19/2026)
  2. DSPy Optimizers (formerly Teleprompters) · DSPy (Stanford NLP), documentation source (accessed 09/19/2026)
  3. Improve your prompts in the developer console · Anthropic, 10/14/2024 (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page