# Prompt engineering

_Level 01 · Direct prompting · sourced_

Writing instructions that get consistent results.


## Try this in a recipe
- [Turn an invoice into a checked record](/gradient_ascent/recipes/invoice-matching.md): Extract a useful JSON record, preserve missing fields, and catch a total that does not reconcile.

## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Compare how instructions shape a response to the same underlying task. Follow the constraints from the English request into the output, then examine what happens when they compete.

**Assumptions:** The examples assume the needed facts are available. A clearer prompt cannot supply missing evidence or grant access to a source.

**Design choices:** Separate essential facts and success criteria from style preferences. Prioritize conflicting requirements rather than accumulating increasingly long instructions.

**Request:** Write a welcoming workshop invitation under 45 words using only the facts below.

**Starting evidence:** Facts: Oak Hall, Saturday 10 am, free, one item per person. Audience: first-time visitors.

**Action and control:** Specify audience, length, tone, and factual boundaries; evaluate the draft against those constraints.

**Stage records (authored, not executed):**

### Brief · v1

Facts = Oak Hall; Saturday 10 am; free; one item per person.
Audience = first-time visitors.
Open question = repair success is not promised.

What changed: The request is separated into supplied facts and an unsupported promise to watch for.

### Prompt · v2

Write a welcoming invitation under 45 words.
Use only these facts: Oak Hall, Saturday 10 am, free, one item per person.
Do not add a repair guarantee.

What changed: The prompt makes audience, length, and factual boundaries explicit; it does not guarantee compliance.

### Candidate comparison

A: Expert repairs guaranteed at our free workshop!
B: New to repair? Join us at Oak Hall this Saturday at 10 am. Admission is free; bring one item and we will explore how to fix it together.
Difference: B removes the unsupported guarantee and restores the supplied details.

What changed: Two authored outputs make the factual difference inspectable. They are not measured responses to a prompt experiment.

### Selected draft · B

New to repair? Join us at Oak Hall this Saturday at 10 am. Admission is free; bring one item and we will explore how to fix it together.

What changed: The selected draft is ready for a human to check; nothing has been sent.

### Review sheet

Facts: all four supplied details retained.
Length: below 45 words.
Promise: no repair guarantee.
Audience fit: editorial judgment, not an objective pass.
Generalization: untested on other events.

What changed: A format check and an editorial judgment are separate kinds of evidence.

### Reusable brief

Replace: event, audience, facts, and desired length.
Keep: separate facts from promises.
Test next: a brief with an unknown venue and a conflicting date.
Accept when: important facts remain correct and the wording fits the audience.

What changed: The reusable output is a brief and review method, not a magic prompt.

**Sample result:** New to repair? Join us at Oak Hall this Saturday at 10 am. Admission is free; bring one item and we will explore how to fix it together.

**Change something — Remove the factual boundary:** Counterexample draft adds: Expert repairs guaranteed. That promise is unsupported even if tone and length fit.

**Decision:** Which check matters beyond tone and length?

**Answer:** Check factual promises against the brief.

**Why:** Change one instruction at a time; stronger wording cannot supply absent facts or guarantee compliance.

**Review criteria:** A rubric comparing factual fidelity, audience fit, and constraints across clearly labeled sample outputs.

**Recovery:** Revise the specific instruction responsible for the failure, then try it on a different input. One polished response is weak evidence of a reusable prompt.

**Adapt it:** Substitute your audience, source material, and output purpose. Keep only constraints that make the result more useful; a word limit or exact format is a local choice.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Compare how instructions shape a response to the same underlying task. Follow the constraints from the English request into the output, then examine what happens when they compete.

**Assumptions:** The examples assume the needed facts are available. A clearer prompt cannot supply missing evidence or grant access to a source.

**Design choices:** Separate essential facts and success criteria from style preferences. Prioritize conflicting requirements rather than accumulating increasingly long instructions.

**Request:** Draft a test procedure using only the approved limits and named framework functions.

**Starting evidence:** DUT brief: log supply voltage; limit not approved. Framework has read_voltage and export_csv.

**Action and control:** Specify format, reuse requirements, and missing-data behavior instead of asking for an unconstrained test plan.

**Stage records (authored, not executed):**

### Input record

DUT brief: log supply voltage; limit not approved. Framework has read_voltage and export_csv.

What changed: Establish the facts supplied for this version of the task.

### Design note

Separate essential facts and success criteria from style preferences. Prioritize conflicting requirements rather than accumulating increasingly long instructions.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Specify format, reuse requirements, and missing-data behavior instead of asking for an unconstrained test plan.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Draft calls for read_voltage and export_csv. Pass/fail limit remains unresolved; ask the engineer before adding it.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Verify every function and limit against the supplied references.

If the result falls short:
Revise the specific instruction responsible for the failure, then try it on a different input. One polished response is weak evidence of a reusable prompt.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Substitute your audience, source material, and output purpose. Keep only constraints that make the result more useful; a word limit or exact format is a local choice.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Draft calls for read_voltage and export_csv. Pass/fail limit remains unresolved; ask the engineer before adding it.

**Change something — Demand a complete procedure with no blanks:** Counterexample invents a 5 V threshold. A complete-looking plan can violate the evidence boundary.

**Decision:** Should a completeness instruction justify inventing a limit?

**Answer:** No; clarify the unapproved limit.

**Why:** Prompt constraints can clarify behavior but cannot create engineering requirements.

**Review criteria:** Verify every function and limit against the supplied references.

**Recovery:** Revise the specific instruction responsible for the failure, then try it on a different input. One polished response is weak evidence of a reusable prompt.

**Adapt it:** Substitute your audience, source material, and output purpose. Keep only constraints that make the result more useful; a word limit or exact format is a local choice.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Compare how instructions shape a response to the same underlying task. Follow the constraints from the English request into the output, then examine what happens when they compete.

**Assumptions:** The examples assume the needed facts are available. A clearer prompt cannot supply missing evidence or grant access to a source.

**Design choices:** Separate essential facts and success criteria from style preferences. Prioritize conflicting requirements rather than accumulating increasingly long instructions.

**Request:** Draft a client update from these notes, separating confirmed dates from estimates.

**Starting evidence:** Migration completed. Training may happen Thursday. No owner has confirmed the training date.

**Action and control:** Ask for sections covering completed work, tentative plans, and decisions needed.

**Stage records (authored, not executed):**

### Input record

Migration completed. Training may happen Thursday. No owner has confirmed the training date.

What changed: Establish the facts supplied for this version of the task.

### Design note

Separate essential facts and success criteria from style preferences. Prioritize conflicting requirements rather than accumulating increasingly long instructions.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Ask for sections covering completed work, tentative plans, and decisions needed.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Migration is complete. Thursday training is tentative; confirmation is pending.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Compare each dated statement with the original notes and its certainty.

If the result falls short:
Revise the specific instruction responsible for the failure, then try it on a different input. One polished response is weak evidence of a reusable prompt.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Substitute your audience, source material, and output purpose. Keep only constraints that make the result more useful; a word limit or exact format is a local choice.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Migration is complete. Thursday training is tentative; confirmation is pending.

**Change something — Request a more confident tone:** Confidence must not turn Thursday into a commitment. Preserve tentative status even in polished wording.

**Decision:** Can a tone instruction promote an estimate to a commitment?

**Answer:** No; tone must preserve factual status.

**Why:** Style constraints and evidence constraints serve different purposes.

**Review criteria:** Compare each dated statement with the original notes and its certainty.

**Recovery:** Revise the specific instruction responsible for the failure, then try it on a different input. One polished response is weak evidence of a reusable prompt.

**Adapt it:** Substitute your audience, source material, and output purpose. Keep only constraints that make the result more useful; a word limit or exact format is a local choice.

Prompt engineering is writing the request itself well: instructions, examples, an assigned role, a
required format, and sometimes an explicit ask to reason before answering. It does not change what
level a technique sits at (this is still level 1, one call, decided entirely by code). It changes
what happens inside that one call, by giving the model more to work with than the bare question
alone.

OpenAI, Anthropic and Google each publish their own guidance for this, and the moves they have in
common (structure, examples, a stated format) carry across models even though the details of how
much each one helps do not[1][2][3]. None of it is worth doing once and
trusting forever: write down what a correct answer looks like before changing a prompt, then check
the new version against the same cases the old one had to pass. This page's example makes that
concrete. It uses the same question and the same source text, asked two ways, checked by a regular
expression rather than by reading the reply and deciding it looks better.

This page is sourced, not measured: every instruction below is checked against a maker's own
prompting guide, and no wording here has been scored against another. It is illustrated.

_The web page for this technique includes an interactive step-through of Level 1 · Prompt engineering. The same steps are described in the sections below._

## Practical guidance

This works in the same chat box you already use, no different interface required. Three moves
carry most of the weight, and you can stack them in a single message.

State the role, then the exact instruction, then the format, in that order. Try: "You are a
parts-desk assistant. List every part number mentioned below, one per line, with its price if the
text states one. If a price isn't given, write 'not stated.'" That's a role, an instruction and a
format in three sentences, and each one narrows what the model can plausibly answer with.

For anything where the shape of the answer matters as much as its content, show one example
instead of only describing it: OpenAI's guide calls one well-chosen example few-shot
learning[1], and Google's guide goes further, recommending that few-shot examples always
be included rather than left out[3]. If you want the reply as a table, a bulleted list,
or a single sentence, say so directly; Google's own guide gives that example verbatim: ask for a
response "as a table, bulleted list, elevator pitch, keywords, sentence, or paragraph"[3],
and that is what comes back.

You'll know it worked when the reply is something you, or a script, can check without rereading a
paragraph: a table with the right number of rows, a list that starts where you asked it to. You'll
know it failed when the model answers the right question in the wrong shape, which usually means
the instruction and the format got buried in the same sentence; put the format on its own line and
ask again.

None of this is worth doing for a question you'll ask once. It earns its keep once you're asking a
close variant of the same thing repeatedly, and there Anthropic's guide has the right frame:
arrive with a clear definition of what a correct answer looks like and a few cases to check a new
version against, and set both up before touching the prompt itself[2]. A prompt that
"reads better" once, on one try, is a different claim from one that holds up on cases you didn't
tune it on.

## Implementation details

The example runs the exact contrast above. It sends the same question, against the same source passage,
to the model two ways. The structured prompt adds a role, an explicit two-field output
format, and one worked example: the moves OpenAI's and Anthropic's guides both
describe[1][2]. The bare prompt is the question and the passage with nothing
else added.

`examples/prompt_engineering/run.py` (lines 35-68)

```python
def run(question: str, model: Model, tracer: Tracer, *, structured: bool = True) -> Answer:
    if structured:
        messages = [
            Message(role="system", content=STRUCTURED_SYSTEM),
            Message(role="user", content=f"{PASSAGE}\n\nQuestion: {question}"),
        ]
        tracer.record(
            kind="code",
            decided_by="code",
            title="Build the structured prompt",
            detail="role + output format + one worked example",
        )
    else:
        messages = [Message(role="user", content=f"{PASSAGE}\n\nQuestion: {question}")]
        tracer.record(kind="code", decided_by="code", title="Build the bare prompt", detail=question)
    completion = model.complete(messages, max_tokens=200)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask the model",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    match = FIELD_RE.search(completion.text)
    citations = ["dw480-manual#8", "parts-list#2"] if match else []
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check the reply against the expected format",
        detail="matched PART/PRICE" if match else "did not match the expected format",
    )
    return Answer(text=completion.text, citations=citations)
```

The check at the end is what turns "did structure help" into something answerable rather than a
matter of taste: `FIELD_RE` looks for exactly `PART: <something> PRICE: <something>` in the reply.
Against this site's `StubModel`, the two replies are scripted by hand for the test that exercises
this example: a well-formatted two-line answer for the structured run, and a hedging sentence
that never states a part number in that shape for the bare one. That is an honest limit on what a
stub run can show: it demonstrates the check a real prompt-tuning loop is built from, not a
finding about how any real model responds to more or less structure. The citations above are the
actual claims about real models; this example is the machinery for testing your own.

The same moves have an engineering use where a model may not invent a number.
[Turning a measurement session into a report](/gradient_ascent/recipes/measurement-writeup/) hands
a model a lab notebook and a table of figures code already computed, inside a system prompt that
spells out the rules: copy every number character for character, and never soften a "cannot say"
verdict into a pass. That is instructions and format at work, in an engineering-test and a
precise-measurement setting alike, while a check confirms the draft added no number of its own.

Run it yourself:

`examples/prompt_engineering/README.md` (lines 16-17)

```text
python -m examples.prompt_engineering --model stub:scripted --structured
python -m examples.prompt_engineering --model stub:scripted --no-structured
```

Every step is `decided_by: "code"`, the same as chat: the code always builds whichever prompt
style it was asked for and asks the model exactly once. What changes between the two runs is
entirely inside the prompt, not in the control flow around it, which is the technique in one
sentence: something you do to the request, not a different level.

## When you do not need this

Skip adding structure if the bare question already gets the right answer every time you try it.
Structure has its own cost: more tokens on every call, and a format instruction that itself needs
testing. And if the format absolutely must be valid on every single call rather than usually
valid, use [structured output](/gradient_ascent/techniques/structured-output/) instead of an
instruction alone: a schema is enforced by the API, a format instruction is only a strong
suggestion the model can still miss.

## Failure modes

### The format holds on easy questions and slips on hard ones

- **How to notice it:** Short, simple questions come back in the requested format every time, but a longer or more unusual question makes the model drop it.
- **How to test for it:** Run the same prompt over the site's harder eval questions (multi-hop, conflicting sources) and score format compliance separately from correctness.

### Instructions that quietly conflict

- **How to notice it:** Two rules in the same prompt pull in different directions, and the model resolves the conflict by picking one without telling you it had to choose.
- **How to test for it:** Read the prompt as a checklist and try to follow it yourself, line by line, as if you were the model given exactly that text and nothing else.

### One example teaches the wrong lesson

- **How to notice it:** The model copies an incidental detail of the worked example (its exact wording, its specific numbers) instead of the pattern the example was meant to show.
- **How to test for it:** Change the specific values in the worked example and ask a new question; check whether the answer stays correct or drifts toward the example's own numbers.

### Tuned on too few cases

- **How to notice it:** The prompt looks great on the handful of questions used to write it and gets measurably worse on questions it never saw while being tuned.
- **How to test for it:** Hold out part of the test set while writing the prompt, then score the finished prompt on the held-out part before it ships.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one question:** 1
- **Tokens in, bare prompt:** ~25
- **Tokens in, structured prompt:** ~140
- **Wall time:** ~0.4s

**Compared with structured output (level 1).** Structured output enforces a schema at the API level instead of asking for a format in words, for a similar token cost: the difference is whether an invalid reply is even possible, not how much it costs to ask.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

Prompt engineering is not its own row in the site's eval; it is a way of improving the score at
whichever level you are already using, by holding the retrieval, the tools and the level fixed and
comparing prompt versions against the same 60-question set (`docs/EVALS.md`). The right test is
A/B, not before/after: run the old prompt and the new prompt over the same questions and compare
scores, since a single "it reads better now" impression on a handful of examples is exactly the
failure the iterating-against-test-cases move above exists to catch.

`scripts/eval_run.py` will not score this example against that set: it sends one fixed passage
to the model two ways and never reads the documents the questions are about, so asking the
runner for a score prints that reason and stops (`docs/EVALS.md`). What to measure for a prompt
change is the pair, old prompt against new, on the same inputs: format adherence, the share of
replies that come back in the shape you asked for, and accuracy of what is in them.

## Run it

**What to monitor.** How often the reply matches the required format, tracked separately from whether the content is correct. A reply can be well-formatted and wrong, or correct and unusable because it broke the format a downstream parser expects.

**Cost at volume.** A longer, more structured prompt costs more input tokens on every call, paid on every request regardless of whether that question needed the structure. A structure that only helps on hard questions is often worth adding conditionally rather than to every prompt.

**How it fails in production.** A prompt tuned against a handful of examples during development meets a wider range of real questions in production, and the format-compliance rate drops because the tuning set didn't cover the phrasing that actually shows up.

**What to log.** The full rendered prompt, not just the template; the raw reply; and whether it matched the expected format, so a format failure in production can be replayed against prompt changes before they ship.

## Try it

1. **Use it.** Take a chat app request you make often and add one instruction, one example of the output you want, and a specific format. Compare the reply against what the plain version gave you.
2. **Build it.** Run both commands from examples/prompt_engineering/README.md with --model stub, then edit STRUCTURED_SYSTEM in run.py to remove the worked example and see whether the test file still passes.
3. **Either lane.** Before you try it, write down exactly what a correct answer to your own question would contain. That written-down version is the test case the iterating-against-test-cases move above depends on.


## Sources

1. [Prompt engineering](https://developers.openai.com/api/docs/guides/prompt-engineering) — OpenAI (accessed 2026-09-19)
2. [Prompt engineering overview](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/overview) — Anthropic (accessed 2026-09-19)
3. [Prompt design strategies](https://ai.google.dev/gemini-api/docs/prompting-strategies) — Google (accessed 2026-09-19)


Last reviewed 2026-09-19.
