Level 01 · Direct prompting

Structured output

Getting answers in a fixed format such as JSON.

Sourced

Concept at a glance

Give the answer a shape your code can read.

SequenceConceptual illustration
Give the answer a shape your code can read.Text + schema leads to Model. Model leads to Validate. A fixed format is useful only when you also validate the returned values.Text + schemaWhat to extract and howModelReturns a structured recordValidateAccept or request a retryGive the answer a shape your code can read.Text + schema leads to Model. Model leads to Validate. A fixed format is useful only when you also validate the returned values.Text + schemaWhat to extract and howModelReturns a structured recordValidateAccept or request a retry
Read the connections in words
  • Text + schema → Model: Returns a structured record.
  • Model → Validate: Accept or request a retry.
Key idea

A fixed format is useful only when you also validate the returned values.

CHOOSE YOUR PERSPECTIVE

Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Structured output: see it in practice.

Constraining an answer to a defined data structure so software can validate and consume it.

What you’ll walk through

Follow an unstructured message into fields another system can use. Inspect both whether the result fits the format and whether each value is supported by the source.

The task in this version

Turn this event email into registration fields. Mark missing facts as unknown.

What you’ll learn to check

Show the friendly form first, optional JSON second, schema validation, and field-by-field source evidence.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Everyday lifeAn authored case with its own evidence, changed condition, and decision.
The task in this example

Turn this event email into registration fields. Mark missing facts as unknown.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Email: Meet at Oak Hall on Saturday. Admission is free. No calendar date or accessibility details supplied.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

The receiving application needs a defined contract, including optional fields, units, and how unknowns are represented.

1 / 6

Apply this to your project

Describe your task to your own model and use Structured output as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Structured output means the reply comes back in a fixed shape (a JSON object with named fields), not a paragraph your code has to parse by guessing. Makers reach it two ways. A JSON-Schema response format constrains which tokens the model may produce next, so OpenAI says a model given one “will always generate responses that adhere to” it, and lists among the benefits “No need to validate or retry incorrectly formatted responses”[1]; Gemini takes the same approach through a schema in response_format[2]. Anthropic supports schema-constrained JSON responses through output_config.format, and separately supports strict tool inputs through strict: true. These can be used independently or together[3]. A direct structured reply without tool execution sits at level 1, one request and one response, in a shape your code chose first.

OpenAI also lists cases where a reply still may not match: a refusal, or a response cut short by the token limit[1]. And no schema check confirms the values are right. Validating your own side and retrying once is still worth doing, the way Pydantic raises “an error with a breakdown of what was wrong”[4].

A model announced in September 2026 takes the idea further: Jev, from TypeSafe AI, generates no text at all, only what TypeSafe calls “typed probabilistic decisions”[5]. It is in early access behind a waitlist, and its published figures are TypeSafe’s own, unmeasured here.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Structured output: extract, validate, retry once

Ask for JSON matching a fixed schema, validate the reply, and retry once if it fails.

Level 1 · Direct prompting
QuestionQuestionFind appliance, load passageFind appliance,load passageMODELAsk for JSONAsk for JSONValidate: one field is wrongValidate: onefield is wrongMODELAsk again with the errorAsk againwith the errorValidate: passesValidate: passesAnswerAnswerQuestionQuestionFind appliance, load passageFind appliance,load passageMODELAsk for JSONAsk for JSONValidate: one field is wrongValidate: onefield is wrongMODELAsk again with the errorAsk againwith the errorValidate: passesValidate: passesAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 06Your code chose

The question arrives

"What is the DW-480's warranty?"
0 tokens · 0 ms

Practical guidance

You’ve used this any time an app turned something you typed or said into a form, a calendar entry, or a spreadsheet row instead of a paragraph of text. Look for the moment right after that: does the app show you the extracted fields on a draft or review screen before it commits to anything, or does it just go and do it? “Add lunch with Sam Thursday at noon” becoming a calendar entry is a model filling in a title, a date and a time; the software worth trusting is the one that shows you those three fields and lets you fix any of them before saving.

If a tool skips that step, or you can’t find a confirmation screen anywhere in it, give it a genuinely ambiguous instruction on purpose: “Set up a payment for the amount in this email,” with no amount stated anywhere, or a date that could mean two different things. Watch what it does with the gap. A well-built feature asks you to confirm the field or fill in the blank yourself; a poorly built one invents something plausible and acts on it, which is how a wrong date or a wrong amount gets through with nobody noticing until later.

Two things are worth checking apart from each other, not as one pass. Did the extraction fill every field it needed, in the right shape, a real date rather than a scrap of text that only looks like one? And separately: is the value actually correct, the date you meant, the amount the document actually states? A tool can pass the first check and fail the second, and a shape that looks valid is not the same claim as a fact that’s true.

None of this needs a second look for something low-stakes you’d catch and fix in five seconds anyway, a draft you were going to reread regardless. It matters for anything costly to get wrong: a payment amount, a shipping address, a date on something legal. That’s where a review step before saving earns its place, and its absence is invisible in a demo, right up until the first time the extraction is wrong.

Implementation details

The example extracts a warranty record for one appliance from evals/corpus/warranty-policy.md: years of full coverage, the years and scope of the limited warranty that follows it, and how many days of coverage apply to commercial or rental use. The schema is five fields, all required.

examples/structured_output/run.py · lines 56–88
def run(question: str, model: Model, tracer: Tracer, *, corpus_dir=DEFAULT_CORPUS_DIR) -> Answer:
    match = APPLIANCE_RE.search(question)
    appliance = match.group(0) if match else "DW-300"
    sections = load_sections(corpus_dir)
    passage = "\n\n".join(sections[cite].text for cite in WARRANTY_SECTIONS)
    tracer.record(kind="code", decided_by="code", title="Find which appliance the question asks about", detail=appliance)
    messages = [
        Message(role="system", content=SYSTEM_PROMPT),
        Message(role="user", content=f"{passage}\n\nAppliance: {appliance}"),
    ]
    record: object = {}
    for attempt in range(MAX_RETRIES + 1):
        completion = model.complete(messages, schema=SCHEMA, max_tokens=200)
        tracer.record(
            kind="model",
            decided_by="code",
            title="Ask the model for JSON" if attempt == 0 else "Ask again with the validation error",
            detail=completion.text[:200],
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
        try:
            record, problems = json.loads(completion.text), None
            problems = _validate(record, appliance)
        except json.JSONDecodeError as exc:
            record, problems = {}, [f"invalid JSON: {exc}"]
        tracer.record(kind="code", decided_by="code", title="Validate against the schema", detail="; ".join(problems) or "valid")
        if not problems:
            return Answer(text=json.dumps(record, sort_keys=True), citations=WARRANTY_SECTIONS)
        if attempt < MAX_RETRIES:
            messages.append(Message(role="user", content=f"That did not validate: {'; '.join(problems)}. Reply again with corrected JSON only."))
    return Answer(text=json.dumps({"error": "did not validate after retry", "last": record}), citations=[])

The retry is deliberately capped at one. _validate checks the reply for every required field, checks that the appliance named in the reply matches the one that was actually asked about (a model can return well-typed JSON about the wrong appliance), and checks that the numeric fields are really integers rather than, say, the string "90 days": a mistake a schema does not always catch, depending on how strictly the backend enforces it. If the first reply fails validation, the code appends the specific problem to the conversation and asks once more; a schema-constrained backend makes the second reply far more likely to be correctly typed, but this example’s own check does not assume that and validates the second reply again regardless. If it is still invalid, the run reports that plainly instead of returning something that never actually passed.

Run it yourself:

examples/structured_output/README.md · lines 15–15
python -m examples.structured_output --model stub:scripted

Every step is decided_by: "code": the schema is fixed, the retry count is fixed, and the model only ever chooses the field values inside whatever shape it was given. Compare this with function calling at level 4, where the model additionally decides whether to use a schema-shaped tool at all: the schema there is the same idea, but the decision of when to reach for it moves from your code to the model.

The same schema-fill pattern serves the bench, too. Reading an instrument’s programming manual and filling one schema row per range and per calibration interval, in ppm of reading and ppm of range with a temperature band and its outside-band coefficient, is the same validate-then-retry extraction as the warranty record above. Code computes the uncertainty budget from the rows; a person checks each row against the manual first, in engineering test and in precise measurement alike, since a right-looking number from the wrong interval or range reads like a correct one.

When you do not need this

Skip the schema, and just read the reply as text, if nothing downstream actually parses it: a chat app showing an answer to a person does not need JSON. And if the field you need is already typed in a fixed, unambiguous format (a form field a person filled in directly, a part number a barcode scanner read), level 0, no model at all reads it directly, with no model and nothing to validate.

Failure modes

Schema-valid, still wrong

How to notice it
Every field is the right type and none are missing, but a value is factually incorrect: the model extracted a real-looking number that is not the one the source actually states.
How to test for it
Compare the extracted values against the source passage by hand on a sample of real runs, not just by checking that the JSON parses.

A model that ignores the schema anyway

How to notice it
Without an API-level guarantee (JSON mode, a forced tool call), the model sometimes wraps the JSON in prose or markdown fences, and a plain parser throws before validation even runs.
How to test for it
Feed the exact raw reply through the same parser production code uses, not a version you cleaned up by hand while debugging.

Retrying on the same mistake

How to notice it
A validation error is sent back and the model makes the same mistake again, because the error message did not actually explain what to change.
How to test for it
Check whether the second reply differs at all from the first; if retries look identical, the retry prompt is not doing its job.

A schema stricter than the task

How to notice it
A field marked required fails validation on a legitimate case where that value genuinely is not knowable (a warranty exclusion with no stated time limit), forcing the model to invent something rather than say so.
How to test for it
Look for retries or failures clustering on one specific kind of input rather than spread evenly across questions.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1–2Model calls, one question
~180Tokens in, first attempt
~40Tokens out
~0.5–0.9sWall time
Compared with chat (level 1)A schema and the fields it requires add a modest number of input tokens over an unstructured reply; the real cost is the retry path, which roughly doubles the call whenever the first reply fails validation.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

Structured output adds a check no plain-text reply can be given: whether the reply is valid against its schema at all, tracked separately from whether the values in it are right. A run can score well on validity and badly on the values, or the reverse, and the two numbers together say more than either alone. Count the retries as a third number. A schema that needs a second attempt on a third of its inputs is a schema to rewrite, not a model to replace.

This example extracts a fixed warranty record rather than answering the question it is handed, so scripts/eval_run.py will not score it against the site’s 60-question set; asking prints that reason and stops (docs/EVALS.md). The set that would measure it is a different one: a list of passages, each with the record it should produce. Score valid-JSON rate, accuracy field by field against those records, and how often the retry was needed.

Run it

What to monitor

Schema-validity rate and semantic-correctness rate, tracked as two separate numbers. A drop in either one means something different and gets fixed differently.

Cost at volume

Roughly one call per extraction, plus a second call for whatever share of replies fail validation the first time. That retry rate is the number to watch, since it is the part of the cost that is not fixed.

How it fails in production

The source text changes shape slightly (a new document template, a field that used to always be present is now sometimes blank) and the extraction starts failing validation at a rate nobody notices until something downstream breaks on missing data.

What to log

The source passage, the full prompt including the schema, every attempt's raw reply, and the validation result for each attempt, so a bad record traces back to which attempt produced it and why.

Try it

  1. Use it

    Find a feature that turns text into a form or a calendar event and give it an ambiguous input. Does it ask you to confirm, or commit to a guess?

  2. Build it

    Run python -m examples.structured_output --model stub:scripted from the repo root: the first record is rejected for a warranty term written as a word; the retry validates. Now rename one field in REQUIRED_FIELDS in examples/structured_output/run.py: both passes fail the same way, and the run ends with an error saying it did not validate after retry, capped at one.

  3. Either lane

    Write the schema you would want for a task in your own life, a recipe or a receipt, before asking a model to fill it. Which fields are truly required, and which would you rather leave blank than guessed?

  4. Build it

    Open the accuracy specs from the manual recipe and find the schema for a row of its table. Which fields would a person have to check against the manual before the row could be trusted?

How it connects

Before, after and instead of this

Move up when

  • Function callingThe model itself has to decide whether to use the schema-shaped output at all, not just fill in values for a shape your code already chose.
Optional: products, tools, and models

8 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 2 more examples
  • Pydantic Tool or framework · Pydantic

    Schema validation

    Checked 09/19/2026
  • Jev Model · TypeSafe AI

    System one decision model

    Checked 09/18/2026
In practice

Extract an invoice record

Ask for supplier, invoice number, date, and amount in a fixed JSON schema, then validate each field.

Out there

Named products, tools and models

Tools7
  • AI SDKVercel · TypeScript AI and agent SDK
  • Claude APIAnthropic · model API
  • Gemini APIGoogle · model API
  • Instructoropen source · structured output library
  • OpenAI APIOpenAI · model API
  • Outlinesdottxt · structured output library
  • PydanticPydantic · schema validation
Models1
  • JevTypeSafe AI · system one decision model

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Structured Outputs · OpenAI (accessed 09/19/2026)
  2. Structured output · Google (accessed 09/19/2026)
  3. Structured outputs · Anthropic (accessed 09/19/2026)
  4. Pydantic Validation · Pydantic (accessed 09/19/2026)
  5. Introducing System One Models & Jev · TypeSafe AI, 09/15/2026 (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page