The example extracts a warranty record for one appliance from
evals/corpus/warranty-policy.md: years of full coverage, the years and scope of the limited
warranty that follows it, and how many days of coverage apply to commercial or rental use. The
schema is five fields, all required.
examples/structured_output/run.py · lines 56–88
def run(question: str, model: Model, tracer: Tracer, *, corpus_dir=DEFAULT_CORPUS_DIR) -> Answer:
match = APPLIANCE_RE.search(question)
appliance = match.group(0) if match else "DW-300"
sections = load_sections(corpus_dir)
passage = "\n\n".join(sections[cite].text for cite in WARRANTY_SECTIONS)
tracer.record(kind="code", decided_by="code", title="Find which appliance the question asks about", detail=appliance)
messages = [
Message(role="system", content=SYSTEM_PROMPT),
Message(role="user", content=f"{passage}\n\nAppliance: {appliance}"),
]
record: object = {}
for attempt in range(MAX_RETRIES + 1):
completion = model.complete(messages, schema=SCHEMA, max_tokens=200)
tracer.record(
kind="model",
decided_by="code",
title="Ask the model for JSON" if attempt == 0 else "Ask again with the validation error",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
try:
record, problems = json.loads(completion.text), None
problems = _validate(record, appliance)
except json.JSONDecodeError as exc:
record, problems = {}, [f"invalid JSON: {exc}"]
tracer.record(kind="code", decided_by="code", title="Validate against the schema", detail="; ".join(problems) or "valid")
if not problems:
return Answer(text=json.dumps(record, sort_keys=True), citations=WARRANTY_SECTIONS)
if attempt < MAX_RETRIES:
messages.append(Message(role="user", content=f"That did not validate: {'; '.join(problems)}. Reply again with corrected JSON only."))
return Answer(text=json.dumps({"error": "did not validate after retry", "last": record}), citations=[])
The retry is deliberately capped at one. _validate checks the reply for every required field,
checks that the appliance named in the reply matches the one that was actually asked about (a
model can return well-typed JSON about the wrong appliance), and checks that the numeric fields
are really integers rather than, say, the string "90 days": a mistake a schema does not always
catch, depending on how strictly the backend enforces it. If the first reply fails validation, the
code appends the specific problem to the conversation and asks once more; a schema-constrained
backend makes the second reply far more likely to be correctly typed, but this example’s own
check does not assume that and validates the second reply again regardless. If it is still
invalid, the run reports that plainly instead of returning something that never actually passed.
Run it yourself:
examples/structured_output/README.md · lines 15–15
python -m examples.structured_output --model stub:scripted
Every step is decided_by: "code": the schema is fixed, the retry count is fixed, and the model
only ever chooses the field values inside whatever shape it was given. Compare this with
function calling at level 4, where the model
additionally decides whether to use a schema-shaped tool at all: the schema there is the same
idea, but the decision of when to reach for it moves from your code to the model.
The same schema-fill pattern serves the bench, too.
Reading an instrument’s programming manual
and filling one schema row per range and per calibration interval, in ppm of reading and
ppm of range with a temperature band and its outside-band coefficient, is the same
validate-then-retry extraction as the warranty record above. Code computes the uncertainty budget
from the rows; a person checks each row against the manual first, in engineering test and in
precise measurement alike, since a right-looking number from the wrong interval or range reads
like a correct one.