# Structured output

_Level 01 · Direct prompting · sourced_

Getting answers in a fixed format such as JSON.


## Try this in a recipe
- [Turn an invoice into a checked record](/gradient_ascent/recipes/invoice-matching.md): Extract a useful JSON record, preserve missing fields, and catch a total that does not reconcile.

## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow an unstructured message into fields another system can use. Inspect both whether the result fits the format and whether each value is supported by the source.

**Assumptions:** The receiving application needs a defined contract, including optional fields, units, and how unknowns are represented.

**Design choices:** Choose a schema around the consumer's needs. Validate shape with code and meaning against evidence; syntactically valid output can still be wrong.

**Request:** Turn this event email into registration fields. Mark missing facts as unknown.

**Starting evidence:** Email: Meet at Oak Hall on Saturday. Admission is free. No calendar date or accessibility details supplied.

**Action and control:** Extract fields, validate their types, then compare each value with its source; structure and truth are separate checks.

**Stage records (authored, not executed):**

### Input record

Email: Meet at Oak Hall on Saturday. Admission is free. No calendar date or accessibility details supplied.

What changed: Establish the facts supplied for this version of the task.

### Design note

Choose a schema around the consumer's needs. Validate shape with code and meaning against evidence; syntactically valid output can still be wrong.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Extract fields, validate their types, then compare each value with its source; structure and truth are separate checks.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Venue: Oak Hall
Day: Saturday
Calendar date: unknown
Price: 0
Accessibility: unknown

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Show the friendly form first, optional JSON second, schema validation, and field-by-field source evidence.

If the result falls short:
Preserve the original input and identify the failed field. Retry a repair for a formatting issue; ask for evidence when the value itself is unknown.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Adapt this to forms, tickets, inventory, or test configurations. Your schema and missing-value policy can differ without changing the distinction between structure and truth.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Venue: Oak Hall
Day: Saturday
Calendar date: unknown
Price: 0
Accessibility: unknown

**Change something — Force all fields to be populated:** A syntactically valid record says accessibility = true and invents a date. Schema validation passes but factual validation fails.

**Decision:** Does valid JSON establish that accessibility is confirmed?

**Answer:** No; the source must support the field.

**Why:** A valid structure may contain wrong facts; missing values need an explicit unknown state instead of invention.

**Review criteria:** Show the friendly form first, optional JSON second, schema validation, and field-by-field source evidence.

**Recovery:** Preserve the original input and identify the failed field. Retry a repair for a formatting issue; ask for evidence when the value itself is unknown.

**Adapt it:** Adapt this to forms, tickets, inventory, or test configurations. Your schema and missing-value policy can differ without changing the distinction between structure and truth.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow an unstructured message into fields another system can use. Inspect both whether the result fits the format and whether each value is supported by the source.

**Assumptions:** The receiving application needs a defined contract, including optional fields, units, and how unknowns are represented.

**Design choices:** Choose a schema around the consumer's needs. Validate shape with code and meaning against evidence; syntactically valid output can still be wrong.

**Request:** Extract acceptance criteria into a reviewable test matrix.

**Starting evidence:** Spec: gain 10 ± 0.5 at 1 kHz. No temperature condition supplied.

**Action and control:** Extract parameter, bounds, units, stimulus, and missing conditions; validate types and source support separately.

**Stage records (authored, not executed):**

### Input record

Spec: gain 10 ± 0.5 at 1 kHz. No temperature condition supplied.

What changed: Establish the facts supplied for this version of the task.

### Design note

Choose a schema around the consumer's needs. Validate shape with code and meaning against evidence; syntactically valid output can still be wrong.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Extract parameter, bounds, units, stimulus, and missing conditions; validate types and source support separately.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Gain: 9.5–10.5; stimulus: 1 kHz; temperature: unknown. Matrix is a draft.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Recompute the bounds and trace every field to the specification.

If the result falls short:
Preserve the original input and identify the failed field. Retry a repair for a formatting issue; ask for evidence when the value itself is unknown.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Adapt this to forms, tickets, inventory, or test configurations. Your schema and missing-value policy can differ without changing the distinction between structure and truth.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Gain: 9.5–10.5; stimulus: 1 kHz; temperature: unknown. Matrix is a draft.

**Change something — Infer room temperature from a past test:** The record is structurally complete but adds an unsupported condition. Flag it for review.

**Decision:** Does a valid matrix establish an approved test condition?

**Answer:** No; each condition needs source support.

**Why:** Schema checks catch shape errors, not invented requirements.

**Review criteria:** Recompute the bounds and trace every field to the specification.

**Recovery:** Preserve the original input and identify the failed field. Retry a repair for a formatting issue; ask for evidence when the value itself is unknown.

**Adapt it:** Adapt this to forms, tickets, inventory, or test configurations. Your schema and missing-value policy can differ without changing the distinction between structure and truth.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow an unstructured message into fields another system can use. Inspect both whether the result fits the format and whether each value is supported by the source.

**Assumptions:** The receiving application needs a defined contract, including optional fields, units, and how unknowns are represented.

**Design choices:** Choose a schema around the consumer's needs. Validate shape with code and meaning against evidence; syntactically valid output can still be wrong.

**Request:** Extract action items from these meeting notes into a tracker draft.

**Starting evidence:** Notes: Priya will confirm shipping by Friday. We should consider a dashboard. No owner assigned to dashboard.

**Action and control:** Distinguish an agreed action from an unassigned suggestion.

**Stage records (authored, not executed):**

### Input record

Notes: Priya will confirm shipping by Friday. We should consider a dashboard. No owner assigned to dashboard.

What changed: Establish the facts supplied for this version of the task.

### Design note

Choose a schema around the consumer's needs. Validate shape with code and meaning against evidence; syntactically valid output can still be wrong.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Distinguish an agreed action from an unassigned suggestion.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Action: confirm shipping; owner Priya; due Friday. Dashboard: suggestion, no assigned owner or deadline.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Check commitments, owners, and dates against the notes before writing to the tracker.

If the result falls short:
Preserve the original input and identify the failed field. Retry a repair for a formatting issue; ask for evidence when the value itself is unknown.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Adapt this to forms, tickets, inventory, or test configurations. Your schema and missing-value policy can differ without changing the distinction between structure and truth.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Action: confirm shipping; owner Priya; due Friday. Dashboard: suggestion, no assigned owner or deadline.

**Change something — Require an owner for every sentence:** A fabricated dashboard owner satisfies the schema but misrepresents the meeting.

**Decision:** Should extraction invent an owner to satisfy a required field?

**Answer:** No; mark unassigned or route for clarification.

**Why:** Structured output needs an explicit representation for missing or inapplicable values.

**Review criteria:** Check commitments, owners, and dates against the notes before writing to the tracker.

**Recovery:** Preserve the original input and identify the failed field. Retry a repair for a formatting issue; ask for evidence when the value itself is unknown.

**Adapt it:** Adapt this to forms, tickets, inventory, or test configurations. Your schema and missing-value policy can differ without changing the distinction between structure and truth.

Structured output means the reply comes back in a fixed shape (a JSON object with named fields),
not a paragraph your code has to parse by guessing. Makers reach it two ways. A JSON-Schema
response format constrains which tokens the model may produce next, so OpenAI says a model given
one "will always generate responses that adhere to" it, and lists among the benefits "No need to
validate or retry incorrectly formatted responses"[1]; Gemini takes the same approach
through a schema in `response_format`[2]. Anthropic supports schema-constrained JSON
responses through `output_config.format`, and separately supports strict tool inputs through
`strict: true`. These can be used independently or together[3]. A direct structured reply
without tool execution sits at level 1, one request and one response, in a shape your code chose first.

OpenAI also lists cases where a reply still may not match: a refusal, or a response cut short by
the token limit[1]. And no schema check confirms the values are right. Validating your
own side and retrying once is still worth doing, the way Pydantic raises "an error with a
breakdown of what was wrong"[4].

A model announced in September 2026 takes the idea further: Jev, from TypeSafe AI, generates no
text at all, only what TypeSafe calls "typed probabilistic decisions"[5]. It is in
early access behind a waitlist, and its published figures are TypeSafe's own, unmeasured here.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 1 · Structured output. The same steps are described in the sections below._

## Practical guidance

You've used this any time an app turned something you typed or said into a form, a calendar entry,
or a spreadsheet row instead of a paragraph of text. Look for the moment right after that: does
the app show you the extracted fields on a draft or review screen before it commits to anything,
or does it just go and do it? "Add lunch with Sam Thursday at noon" becoming a calendar entry is a
model filling in a title, a date and a time; the software worth trusting is the one that shows you
those three fields and lets you fix any of them before saving.

If a tool skips that step, or you can't find a confirmation screen anywhere in it, give it a
genuinely ambiguous instruction on purpose: "Set up a payment for the amount in this email," with
no amount stated anywhere, or a date that could mean two different things. Watch what it does with
the gap. A well-built feature asks you to confirm the field or fill in the blank yourself; a
poorly built one invents something plausible and acts on it, which is how a wrong date or a wrong
amount gets through with nobody noticing until later.

Two things are worth checking apart from each other, not as one pass. Did the extraction fill
every field it needed, in the right shape, a real date rather than a scrap of text that only looks
like one? And separately: is the value actually correct, the date you meant, the amount the
document actually states? A tool can pass the first check and fail the second, and a shape that
looks valid is not the same claim as a fact that's true.

None of this needs a second look for something low-stakes you'd catch and fix in five seconds
anyway, a draft you were going to reread regardless. It matters for anything costly to get wrong: a
payment amount, a shipping address, a date on something legal. That's where a review step before
saving earns its place, and its absence is invisible in a demo, right up until the first time the
extraction is wrong.

## Implementation details

The example extracts a warranty record for one appliance from
`evals/corpus/warranty-policy.md`: years of full coverage, the years and scope of the limited
warranty that follows it, and how many days of coverage apply to commercial or rental use. The
schema is five fields, all required.

`examples/structured_output/run.py` (lines 56-88)

```python
def run(question: str, model: Model, tracer: Tracer, *, corpus_dir=DEFAULT_CORPUS_DIR) -> Answer:
    match = APPLIANCE_RE.search(question)
    appliance = match.group(0) if match else "DW-300"
    sections = load_sections(corpus_dir)
    passage = "\n\n".join(sections[cite].text for cite in WARRANTY_SECTIONS)
    tracer.record(kind="code", decided_by="code", title="Find which appliance the question asks about", detail=appliance)
    messages = [
        Message(role="system", content=SYSTEM_PROMPT),
        Message(role="user", content=f"{passage}\n\nAppliance: {appliance}"),
    ]
    record: object = {}
    for attempt in range(MAX_RETRIES + 1):
        completion = model.complete(messages, schema=SCHEMA, max_tokens=200)
        tracer.record(
            kind="model",
            decided_by="code",
            title="Ask the model for JSON" if attempt == 0 else "Ask again with the validation error",
            detail=completion.text[:200],
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
        try:
            record, problems = json.loads(completion.text), None
            problems = _validate(record, appliance)
        except json.JSONDecodeError as exc:
            record, problems = {}, [f"invalid JSON: {exc}"]
        tracer.record(kind="code", decided_by="code", title="Validate against the schema", detail="; ".join(problems) or "valid")
        if not problems:
            return Answer(text=json.dumps(record, sort_keys=True), citations=WARRANTY_SECTIONS)
        if attempt < MAX_RETRIES:
            messages.append(Message(role="user", content=f"That did not validate: {'; '.join(problems)}. Reply again with corrected JSON only."))
    return Answer(text=json.dumps({"error": "did not validate after retry", "last": record}), citations=[])
```

The retry is deliberately capped at one. `_validate` checks the reply for every required field,
checks that the appliance named in the reply matches the one that was actually asked about (a
model can return well-typed JSON about the wrong appliance), and checks that the numeric fields
are really integers rather than, say, the string `"90 days"`: a mistake a schema does not always
catch, depending on how strictly the backend enforces it. If the first reply fails validation, the
code appends the specific problem to the conversation and asks once more; a schema-constrained
backend makes the second reply far more likely to be correctly typed, but this example's own
check does not assume that and validates the second reply again regardless. If it is still
invalid, the run reports that plainly instead of returning something that never actually passed.

Run it yourself:

`examples/structured_output/README.md` (lines 15-15)

```text
python -m examples.structured_output --model stub:scripted
```

Every step is `decided_by: "code"`: the schema is fixed, the retry count is fixed, and the model
only ever chooses the field values inside whatever shape it was given. Compare this with
[function calling](/gradient_ascent/techniques/function-calling/) at level 4, where the model
additionally decides whether to use a schema-shaped tool at all: the schema there is the same
idea, but the decision of when to reach for it moves from your code to the model.

The same schema-fill pattern serves the bench, too.
[Reading an instrument's programming manual
and filling one schema row per range and per calibration interval](/gradient_ascent/recipes/accuracy-specs-from-the-manual/), in ppm of reading and
ppm of range with a temperature band and its outside-band coefficient, is the same
validate-then-retry extraction as the warranty record above. Code computes the uncertainty budget
from the rows; a person checks each row against the manual first, in engineering test and in
precise measurement alike, since a right-looking number from the wrong interval or range reads
like a correct one.

## When you do not need this

Skip the schema, and just read the reply as text, if nothing downstream actually parses it: a
chat app showing an answer to a person does not need JSON. And if the field you need is already
typed in a fixed, unambiguous format (a form field a person filled in directly, a part number a
barcode scanner read), [level 0, no model at all](/gradient_ascent/techniques/order-zero/) reads it directly,
with no model and nothing to validate.

## Failure modes

### Schema-valid, still wrong

- **How to notice it:** Every field is the right type and none are missing, but a value is factually incorrect: the model extracted a real-looking number that is not the one the source actually states.
- **How to test for it:** Compare the extracted values against the source passage by hand on a sample of real runs, not just by checking that the JSON parses.

### A model that ignores the schema anyway

- **How to notice it:** Without an API-level guarantee (JSON mode, a forced tool call), the model sometimes wraps the JSON in prose or markdown fences, and a plain parser throws before validation even runs.
- **How to test for it:** Feed the exact raw reply through the same parser production code uses, not a version you cleaned up by hand while debugging.

### Retrying on the same mistake

- **How to notice it:** A validation error is sent back and the model makes the same mistake again, because the error message did not actually explain what to change.
- **How to test for it:** Check whether the second reply differs at all from the first; if retries look identical, the retry prompt is not doing its job.

### A schema stricter than the task

- **How to notice it:** A field marked required fails validation on a legitimate case where that value genuinely is not knowable (a warranty exclusion with no stated time limit), forcing the model to invent something rather than say so.
- **How to test for it:** Look for retries or failures clustering on one specific kind of input rather than spread evenly across questions.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one question:** 1–2
- **Tokens in, first attempt:** ~180
- **Tokens out:** ~40
- **Wall time:** ~0.5–0.9s

**Compared with chat (level 1).** A schema and the fields it requires add a modest number of input tokens over an unstructured reply; the real cost is the retry path, which roughly doubles the call whenever the first reply fails validation.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

Structured output adds a check no plain-text reply can be given: whether the reply is valid
against its schema at all, tracked separately from whether the values in it are right. A run can
score well on validity and badly on the values, or the reverse, and the two numbers together say
more than either alone. Count the retries as a third number. A schema that needs a second
attempt on a third of its inputs is a schema to rewrite, not a model to replace.

This example extracts a fixed warranty record rather than answering the question it is handed,
so `scripts/eval_run.py` will not score it against the site's 60-question set; asking prints
that reason and stops (`docs/EVALS.md`). The set that would measure it is a different one: a
list of passages, each with the record it should produce. Score valid-JSON rate, accuracy field
by field against those records, and how often the retry was needed.

## Run it

**What to monitor.** Schema-validity rate and semantic-correctness rate, tracked as two separate numbers. A drop in either one means something different and gets fixed differently.

**Cost at volume.** Roughly one call per extraction, plus a second call for whatever share of replies fail validation the first time. That retry rate is the number to watch, since it is the part of the cost that is not fixed.

**How it fails in production.** The source text changes shape slightly (a new document template, a field that used to always be present is now sometimes blank) and the extraction starts failing validation at a rate nobody notices until something downstream breaks on missing data.

**What to log.** The source passage, the full prompt including the schema, every attempt's raw reply, and the validation result for each attempt, so a bad record traces back to which attempt produced it and why.

## Try it

1. **Use it.** Find a feature that turns text into a form or a calendar event and give it an ambiguous input. Does it ask you to confirm, or commit to a guess?
2. **Build it.** Run python -m examples.structured_output --model stub:scripted from the repo root: the first record is rejected for a warranty term written as a word; the retry validates. Now rename one field in REQUIRED_FIELDS in examples/structured_output/run.py: both passes fail the same way, and the run ends with an error saying it did not validate after retry, capped at one.
3. **Either lane.** Write the schema you would want for a task in your own life, a recipe or a receipt, before asking a model to fill it. Which fields are truly required, and which would you rather leave blank than guessed?
4. **Build it.** Open the accuracy specs from the manual recipe and find the schema for a row of its table. Which fields would a person have to check against the manual before the row could be trusted?


## Sources

1. [Structured Outputs](https://developers.openai.com/api/docs/guides/structured-outputs) — OpenAI (accessed 2026-09-19)
2. [Structured output](https://ai.google.dev/gemini-api/docs/structured-output) — Google (accessed 2026-09-19)
3. [Structured outputs](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) — Anthropic (accessed 2026-09-19)
4. [Pydantic Validation](https://pydantic.dev/docs/validation/latest/get-started/) — Pydantic (accessed 2026-09-19)
5. [Introducing System One Models & Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) — TypeSafe AI, 2026-09-15 (accessed 2026-09-19)


Last reviewed 2026-09-19.
