Primary sources
- Prompt engineering · OpenAI (accessed 09/19/2026)
- Prompt engineering overview · Anthropic (accessed 09/19/2026)
- Prompt design strategies · Google (accessed 09/19/2026)
Writing instructions that get consistent results.
Sourced
Concept at a glance
Clear instructions shape the response without changing the model itself.
Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.
Designing instructions, examples, and output expectations to steer a model's response.
Compare how instructions shape a response to the same underlying task. Follow the constraints from the English request into the output, then examine what happens when they compete.
Write a welcoming workshop invitation under 45 words using only the facts below.
A rubric comparing factual fidelity, audience fit, and constraints across clearly labeled sample outputs.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Write a welcoming workshop invitation under 45 words using only the facts below.
Authored case. Select any record below; nothing is sent to a model.What changed: The request is separated into supplied facts and an unsupported promise to watch for.
The examples assume the needed facts are available. A clearer prompt cannot supply missing evidence or grant access to a source.
Describe your task to your own model and use Prompt engineering as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
Prompt engineering is writing the request itself well: instructions, examples, an assigned role, a required format, and sometimes an explicit ask to reason before answering. It does not change what level a technique sits at (this is still level 1, one call, decided entirely by code). It changes what happens inside that one call, by giving the model more to work with than the bare question alone.
OpenAI, Anthropic and Google each publish their own guidance for this, and the moves they have in common (structure, examples, a stated format) carry across models even though the details of how much each one helps do not[1][2][3]. None of it is worth doing once and trusting forever: write down what a correct answer looks like before changing a prompt, then check the new version against the same cases the old one had to pass. This page’s example makes that concrete. It uses the same question and the same source text, asked two ways, checked by a regular expression rather than by reading the reply and deciding it looks better.
This page is sourced, not measured: every instruction below is checked against a maker’s own prompting guide, and no wording here has been scored against another. It is illustrated.
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
Add a role, a fixed format and one worked example to the same question and source passage.
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
"What is the DW-480's drain pump part number and price?" passage: "...HLV-2205, available through authorized service...priced at $52.00..."
This works in the same chat box you already use, no different interface required. Three moves carry most of the weight, and you can stack them in a single message.
State the role, then the exact instruction, then the format, in that order. Try: “You are a parts-desk assistant. List every part number mentioned below, one per line, with its price if the text states one. If a price isn’t given, write ‘not stated.’” That’s a role, an instruction and a format in three sentences, and each one narrows what the model can plausibly answer with.
For anything where the shape of the answer matters as much as its content, show one example instead of only describing it: OpenAI’s guide calls one well-chosen example few-shot learning[1], and Google’s guide goes further, recommending that few-shot examples always be included rather than left out[3]. If you want the reply as a table, a bulleted list, or a single sentence, say so directly; Google’s own guide gives that example verbatim: ask for a response “as a table, bulleted list, elevator pitch, keywords, sentence, or paragraph”[3], and that is what comes back.
You’ll know it worked when the reply is something you, or a script, can check without rereading a paragraph: a table with the right number of rows, a list that starts where you asked it to. You’ll know it failed when the model answers the right question in the wrong shape, which usually means the instruction and the format got buried in the same sentence; put the format on its own line and ask again.
None of this is worth doing for a question you’ll ask once. It earns its keep once you’re asking a close variant of the same thing repeatedly, and there Anthropic’s guide has the right frame: arrive with a clear definition of what a correct answer looks like and a few cases to check a new version against, and set both up before touching the prompt itself[2]. A prompt that “reads better” once, on one try, is a different claim from one that holds up on cases you didn’t tune it on.
The example runs the exact contrast above. It sends the same question, against the same source passage, to the model two ways. The structured prompt adds a role, an explicit two-field output format, and one worked example: the moves OpenAI’s and Anthropic’s guides both describe[1][2]. The bare prompt is the question and the passage with nothing else added.
def run(question: str, model: Model, tracer: Tracer, *, structured: bool = True) -> Answer:
if structured:
messages = [
Message(role="system", content=STRUCTURED_SYSTEM),
Message(role="user", content=f"{PASSAGE}\n\nQuestion: {question}"),
]
tracer.record(
kind="code",
decided_by="code",
title="Build the structured prompt",
detail="role + output format + one worked example",
)
else:
messages = [Message(role="user", content=f"{PASSAGE}\n\nQuestion: {question}")]
tracer.record(kind="code", decided_by="code", title="Build the bare prompt", detail=question)
completion = model.complete(messages, max_tokens=200)
tracer.record(
kind="model",
decided_by="code",
title="Ask the model",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
match = FIELD_RE.search(completion.text)
citations = ["dw480-manual#8", "parts-list#2"] if match else []
tracer.record(
kind="code",
decided_by="code",
title="Check the reply against the expected format",
detail="matched PART/PRICE" if match else "did not match the expected format",
)
return Answer(text=completion.text, citations=citations)The check at the end is what turns “did structure help” into something answerable rather than a
matter of taste: FIELD_RE looks for exactly PART: <something> PRICE: <something> in the reply.
Against this site’s StubModel, the two replies are scripted by hand for the test that exercises
this example: a well-formatted two-line answer for the structured run, and a hedging sentence
that never states a part number in that shape for the bare one. That is an honest limit on what a
stub run can show: it demonstrates the check a real prompt-tuning loop is built from, not a
finding about how any real model responds to more or less structure. The citations above are the
actual claims about real models; this example is the machinery for testing your own.
The same moves have an engineering use where a model may not invent a number. Turning a measurement session into a report hands a model a lab notebook and a table of figures code already computed, inside a system prompt that spells out the rules: copy every number character for character, and never soften a “cannot say” verdict into a pass. That is instructions and format at work, in an engineering-test and a precise-measurement setting alike, while a check confirms the draft added no number of its own.
Run it yourself:
python -m examples.prompt_engineering --model stub:scripted --structured
python -m examples.prompt_engineering --model stub:scripted --no-structuredEvery step is decided_by: "code", the same as chat: the code always builds whichever prompt
style it was asked for and asks the model exactly once. What changes between the two runs is
entirely inside the prompt, not in the control flow around it, which is the technique in one
sentence: something you do to the request, not a different level.
Skip adding structure if the bare question already gets the right answer every time you try it. Structure has its own cost: more tokens on every call, and a format instruction that itself needs testing. And if the format absolutely must be valid on every single call rather than usually valid, use structured output instead of an instruction alone: a schema is enforced by the API, a format instruction is only a strong suggestion the model can still miss.
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
Prompt engineering is not its own row in the site’s eval; it is a way of improving the score at
whichever level you are already using, by holding the retrieval, the tools and the level fixed and
comparing prompt versions against the same 60-question set (docs/EVALS.md). The right test is
A/B, not before/after: run the old prompt and the new prompt over the same questions and compare
scores, since a single “it reads better now” impression on a handful of examples is exactly the
failure the iterating-against-test-cases move above exists to catch.
scripts/eval_run.py will not score this example against that set: it sends one fixed passage
to the model two ways and never reads the documents the questions are about, so asking the
runner for a score prints that reason and stops (docs/EVALS.md). What to measure for a prompt
change is the pair, old prompt against new, on the same inputs: format adherence, the share of
replies that come back in the shape you asked for, and accuracy of what is in them.
How often the reply matches the required format, tracked separately from whether the content is correct. A reply can be well-formatted and wrong, or correct and unusable because it broke the format a downstream parser expects.
A longer, more structured prompt costs more input tokens on every call, paid on every request regardless of whether that question needed the structure. A structure that only helps on hard questions is often worth adding conditionally rather than to every prompt.
A prompt tuned against a handful of examples during development meets a wider range of real questions in production, and the format-compliance rate drops because the tuning set didn't cover the phrasing that actually shows up.
The full rendered prompt, not just the template; the raw reply; and whether it matched the expected format, so a format failure in production can be replayed against prompt changes before they ship.
Take a chat app request you make often and add one instruction, one example of the output you want, and a specific format. Compare the reply against what the plain version gave you.
Run both commands from examples/prompt_engineering/README.md with --model stub, then edit STRUCTURED_SYSTEM in run.py to remove the worked example and see whether the test file still passes.
Before you try it, write down exactly what a correct answer to your own question would contain. That written-down version is the test case the iterating-against-test-cases move above depends on.
5 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.
Prompt engineering guide
Maker’s documentation Checked 09/18/2026Prompt engineering guide
Maker’s documentation Checked 09/18/2026Prompt engineering guide
Maker’s documentation Checked 09/18/2026Prompt generation tool
Maker’s documentation Checked 09/19/2026Prompt optimization tool
Maker’s documentation Checked 09/18/2026Specify the audience, tone, required facts, and an example of the response you want.
Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page