Level 01 · Direct prompting

Prompt engineering

Writing instructions that get consistent results.

Sourced

Concept at a glance

Make the request easier to get right.

SequenceConceptual illustration
Make the request easier to get right.Write the brief leads to Model. Model leads to Check the response. Clear instructions shape the response without changing the model itself.Write the briefGoal, context, examplesModelWorks from your requestCheck the responseAgainst your instructionsMake the request easier to get right.Write the brief leads to Model. Model leads to Check the response. Clear instructions shape the response without changing the model itself.Write the briefGoal, context, examplesModelWorks from your requestCheck the responseAgainst your instructions
Read the connections in words
  • Write the brief → Model: Works from your request.
  • Model → Check the response: Against your instructions.
Key idea

Clear instructions shape the response without changing the model itself.

CHOOSE YOUR PERSPECTIVE

Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Prompt engineering: see it in practice.

Designing instructions, examples, and output expectations to steer a model's response.

What you’ll walk through

Compare how instructions shape a response to the same underlying task. Follow the constraints from the English request into the output, then examine what happens when they compete.

The task in this version

Write a welcoming workshop invitation under 45 words using only the facts below.

What you’ll learn to check

A rubric comparing factual fidelity, audience fit, and constraints across clearly labeled sample outputs.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Everyday lifeAn authored case with its own evidence, changed condition, and decision.
The task in this example

Write a welcoming workshop invitation under 45 words using only the facts below.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Brief · v1
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Facts = Oak Hall; Saturday 10 am; free; one item per person. Audience = first-time visitors. Open question = repair success is not promised.

What changed: The request is separated into supplied facts and an unsupported promise to watch for.

WHY THIS MATTERS

What this case assumes

The examples assume the needed facts are available. A clearer prompt cannot supply missing evidence or grant access to a source.

1 / 6

Apply this to your project

Describe your task to your own model and use Prompt engineering as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Prompt engineering is writing the request itself well: instructions, examples, an assigned role, a required format, and sometimes an explicit ask to reason before answering. It does not change what level a technique sits at (this is still level 1, one call, decided entirely by code). It changes what happens inside that one call, by giving the model more to work with than the bare question alone.

OpenAI, Anthropic and Google each publish their own guidance for this, and the moves they have in common (structure, examples, a stated format) carry across models even though the details of how much each one helps do not[1][2][3]. None of it is worth doing once and trusting forever: write down what a correct answer looks like before changing a prompt, then check the new version against the same cases the old one had to pass. This page’s example makes that concrete. It uses the same question and the same source text, asked two ways, checked by a regular expression rather than by reading the reply and deciding it looks better.

This page is sourced, not measured: every instruction below is checked against a maker’s own prompting guide, and no wording here has been scored against another. It is illustrated.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Prompt engineering: same facts, a structured ask

Add a role, a fixed format and one worked example to the same question and source passage.

Level 1 · Direct prompting
Question + source passageQuestion +source passageAdd role, format, one exampleAdd role, format,one exampleMODELanswers onceanswers onceCheck reply against the formatCheck replyagainst the formatAnswerAnswerQuestion + source passageQuestion +source passageAdd role, format, one exampleAdd role, format,one exampleMODELanswers onceanswers onceCheck reply against the formatCheck replyagainst the formatAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 04Your code chose

The question and a source passage arrive

"What is the DW-480's drain pump part number and price?"
passage: "...HLV-2205, available through authorized
service...priced at $52.00..."
0 tokens · 0 ms

Practical guidance

This works in the same chat box you already use, no different interface required. Three moves carry most of the weight, and you can stack them in a single message.

State the role, then the exact instruction, then the format, in that order. Try: “You are a parts-desk assistant. List every part number mentioned below, one per line, with its price if the text states one. If a price isn’t given, write ‘not stated.’” That’s a role, an instruction and a format in three sentences, and each one narrows what the model can plausibly answer with.

For anything where the shape of the answer matters as much as its content, show one example instead of only describing it: OpenAI’s guide calls one well-chosen example few-shot learning[1], and Google’s guide goes further, recommending that few-shot examples always be included rather than left out[3]. If you want the reply as a table, a bulleted list, or a single sentence, say so directly; Google’s own guide gives that example verbatim: ask for a response “as a table, bulleted list, elevator pitch, keywords, sentence, or paragraph”[3], and that is what comes back.

You’ll know it worked when the reply is something you, or a script, can check without rereading a paragraph: a table with the right number of rows, a list that starts where you asked it to. You’ll know it failed when the model answers the right question in the wrong shape, which usually means the instruction and the format got buried in the same sentence; put the format on its own line and ask again.

None of this is worth doing for a question you’ll ask once. It earns its keep once you’re asking a close variant of the same thing repeatedly, and there Anthropic’s guide has the right frame: arrive with a clear definition of what a correct answer looks like and a few cases to check a new version against, and set both up before touching the prompt itself[2]. A prompt that “reads better” once, on one try, is a different claim from one that holds up on cases you didn’t tune it on.

Implementation details

The example runs the exact contrast above. It sends the same question, against the same source passage, to the model two ways. The structured prompt adds a role, an explicit two-field output format, and one worked example: the moves OpenAI’s and Anthropic’s guides both describe[1][2]. The bare prompt is the question and the passage with nothing else added.

examples/prompt_engineering/run.py · lines 35–68
def run(question: str, model: Model, tracer: Tracer, *, structured: bool = True) -> Answer:
    if structured:
        messages = [
            Message(role="system", content=STRUCTURED_SYSTEM),
            Message(role="user", content=f"{PASSAGE}\n\nQuestion: {question}"),
        ]
        tracer.record(
            kind="code",
            decided_by="code",
            title="Build the structured prompt",
            detail="role + output format + one worked example",
        )
    else:
        messages = [Message(role="user", content=f"{PASSAGE}\n\nQuestion: {question}")]
        tracer.record(kind="code", decided_by="code", title="Build the bare prompt", detail=question)
    completion = model.complete(messages, max_tokens=200)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask the model",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    match = FIELD_RE.search(completion.text)
    citations = ["dw480-manual#8", "parts-list#2"] if match else []
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check the reply against the expected format",
        detail="matched PART/PRICE" if match else "did not match the expected format",
    )
    return Answer(text=completion.text, citations=citations)

The check at the end is what turns “did structure help” into something answerable rather than a matter of taste: FIELD_RE looks for exactly PART: <something> PRICE: <something> in the reply. Against this site’s StubModel, the two replies are scripted by hand for the test that exercises this example: a well-formatted two-line answer for the structured run, and a hedging sentence that never states a part number in that shape for the bare one. That is an honest limit on what a stub run can show: it demonstrates the check a real prompt-tuning loop is built from, not a finding about how any real model responds to more or less structure. The citations above are the actual claims about real models; this example is the machinery for testing your own.

The same moves have an engineering use where a model may not invent a number. Turning a measurement session into a report hands a model a lab notebook and a table of figures code already computed, inside a system prompt that spells out the rules: copy every number character for character, and never soften a “cannot say” verdict into a pass. That is instructions and format at work, in an engineering-test and a precise-measurement setting alike, while a check confirms the draft added no number of its own.

Run it yourself:

examples/prompt_engineering/README.md · lines 16–17
python -m examples.prompt_engineering --model stub:scripted --structured
python -m examples.prompt_engineering --model stub:scripted --no-structured

Every step is decided_by: "code", the same as chat: the code always builds whichever prompt style it was asked for and asks the model exactly once. What changes between the two runs is entirely inside the prompt, not in the control flow around it, which is the technique in one sentence: something you do to the request, not a different level.

When you do not need this

Skip adding structure if the bare question already gets the right answer every time you try it. Structure has its own cost: more tokens on every call, and a format instruction that itself needs testing. And if the format absolutely must be valid on every single call rather than usually valid, use structured output instead of an instruction alone: a schema is enforced by the API, a format instruction is only a strong suggestion the model can still miss.

Failure modes

The format holds on easy questions and slips on hard ones

How to notice it
Short, simple questions come back in the requested format every time, but a longer or more unusual question makes the model drop it.
How to test for it
Run the same prompt over the site's harder eval questions (multi-hop, conflicting sources) and score format compliance separately from correctness.

Instructions that quietly conflict

How to notice it
Two rules in the same prompt pull in different directions, and the model resolves the conflict by picking one without telling you it had to choose.
How to test for it
Read the prompt as a checklist and try to follow it yourself, line by line, as if you were the model given exactly that text and nothing else.

One example teaches the wrong lesson

How to notice it
The model copies an incidental detail of the worked example (its exact wording, its specific numbers) instead of the pattern the example was meant to show.
How to test for it
Change the specific values in the worked example and ask a new question; check whether the answer stays correct or drifts toward the example's own numbers.

Tuned on too few cases

How to notice it
The prompt looks great on the handful of questions used to write it and gets measurably worse on questions it never saw while being tuned.
How to test for it
Hold out part of the test set while writing the prompt, then score the finished prompt on the held-out part before it ships.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1Model calls, one question
~25Tokens in, bare prompt
~140Tokens in, structured prompt
~0.4sWall time
Compared with structured output (level 1)Structured output enforces a schema at the API level instead of asking for a format in words, for a similar token cost: the difference is whether an invalid reply is even possible, not how much it costs to ask.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

Prompt engineering is not its own row in the site’s eval; it is a way of improving the score at whichever level you are already using, by holding the retrieval, the tools and the level fixed and comparing prompt versions against the same 60-question set (docs/EVALS.md). The right test is A/B, not before/after: run the old prompt and the new prompt over the same questions and compare scores, since a single “it reads better now” impression on a handful of examples is exactly the failure the iterating-against-test-cases move above exists to catch.

scripts/eval_run.py will not score this example against that set: it sends one fixed passage to the model two ways and never reads the documents the questions are about, so asking the runner for a score prints that reason and stops (docs/EVALS.md). What to measure for a prompt change is the pair, old prompt against new, on the same inputs: format adherence, the share of replies that come back in the shape you asked for, and accuracy of what is in them.

Run it

What to monitor

How often the reply matches the required format, tracked separately from whether the content is correct. A reply can be well-formatted and wrong, or correct and unusable because it broke the format a downstream parser expects.

Cost at volume

A longer, more structured prompt costs more input tokens on every call, paid on every request regardless of whether that question needed the structure. A structure that only helps on hard questions is often worth adding conditionally rather than to every prompt.

How it fails in production

A prompt tuned against a handful of examples during development meets a wider range of real questions in production, and the format-compliance rate drops because the tuning set didn't cover the phrasing that actually shows up.

What to log

The full rendered prompt, not just the template; the raw reply; and whether it matched the expected format, so a format failure in production can be replayed against prompt changes before they ship.

Try it

  1. Use it

    Take a chat app request you make often and add one instruction, one example of the output you want, and a specific format. Compare the reply against what the plain version gave you.

  2. Build it

    Run both commands from examples/prompt_engineering/README.md with --model stub, then edit STRUCTURED_SYSTEM in run.py to remove the worked example and see whether the test file still passes.

  3. Either lane

    Before you try it, write down exactly what a correct answer to your own question would contain. That written-down version is the test case the iterating-against-test-cases move above depends on.

How it connects

Before, after and instead of this

Read first

Move up when

  • Structured outputThe reply's format has to be valid on every single call, not just usually valid.

Instead of

Optional: products, tools, and models

5 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Write a repeatable support reply

Specify the audience, tone, required facts, and an example of the response you want.

Out there

Named products, tools and models

Tools5
  • Prompt Design StrategiesGoogle · prompt engineering guide
  • Prompt EngineeringOpenAI · prompt engineering guide
  • Prompt Engineering OverviewAnthropic · prompt engineering guide
  • Prompt GeneratorAnthropic · prompt generation tool
  • Prompt ImproverAnthropic · prompt optimization tool

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Prompt engineering · OpenAI (accessed 09/19/2026)
  2. Prompt engineering overview · Anthropic (accessed 09/19/2026)
  3. Prompt design strategies · Google (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page