Recipe

Sort an inbox

Sorts mail into fixed categories and produces structured output. A person approves anything that gets sent. The categories are known in advance, so an agent is not needed.

SourcedNeeds level 3

A small company’s shared inbox gets everything: billing questions, product-support requests, general questions, and the occasional message that fits none of those. Someone has to read each one, decide where it goes, and get the details into the system that handles that category, instead of leaving it as an email somebody must open and reread field by field. Nobody wants a refund request quietly filed into a queue with no person ever seeing the dollar figure in it.

Inbox triage does that in three fixed steps: classify the message into one of a known, small set of categories, turn it into a structured record, and file it to the matching queue. Anything the classifier isn’t confident about, or that names a dollar amount, waits for a person first.

Example run

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Sort an inbox, assembled

Classify the message, extract a ticket record, and pause before filing anything costly or unclear.

Level 3 · Workflows
Message arrivesMessage arrivesMODELClassify the categoryClassifythe categoryRoute to the billing handlerRoute to thebilling handlerMODELExtract ticket fieldsExtractticket fieldsCheck the gateCheck the gatePERSONPerson decidesPerson decidesFile or send backFile or send backTicket filedTicket filed
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 07Your code chose

The message arrives

"Subject: Refund? — the drain pump I bought (HLV-2205)
failed after three weeks, can I get the $46.00 back?"
0 tokens · 0 ms

Walkthrough

The three steps compose the runnable code already on each technique’s own page, unmodified in shape: examples/routing/run.py’s classify-then-dispatch pattern, examples/structured_output/ run.py’s ask-validate-retry-once pattern, and examples/human_in_the_loop/run.py’s check-thresholds-and-resume pattern. All three are written against the site’s shared document set, not a literal inbox, so this exact version is not in the repository. Composing them means keeping the same functions and swapping in this job’s own labels (billing, support, general, unclear instead of lookup, numeric, unclear) and its own schema, a ticket record instead of a warranty record. The parsing, the routing table and the one-retry contract carry over unchanged; none of them inspect what the labels are called.

The run above follows one message: a refund request naming a part and a dollar amount. It classifies as billing, and the handler extracts a ticket record (category, summary, amount, priority), the same shape structured output’s own example validates before accepting. The gate trips on high_cost, since the record names a dollar figure, and the run pauses rather than filing it. A person sees the message and the record, approves it, and the code files the ticket. A message that classifies unclear pauses by the same code path for the opposite reason: the fallback found nothing confident enough to hand a handler at all.

Two things to settle before this runs on real mail. Every message body goes to whichever model classifies it, so where that model runs is a decision about customer data and not only about accuracy. And log the raw classification text, the parsed label, the record, the gate’s reason and the person’s decision together: an approval with no record of what the reviewer was shown cannot be audited afterwards.

What to measure

Build a labeled set for this job: fifty to a hundred real or synthetic messages, hand-labeled with the category and the ticket fields a person would write down. Three numbers, each testing a different piece: routing accuracy (does the label match the human one); extraction accuracy field by field, with the valid-JSON rate and retry count tracked separately; and the pause rate split by reason, checked against a sample of messages that did not pause, to see whether any should have. No result file exists for routing, structured output or human approval yet (see docs/EVALS.md), so this recipe claims no score.

Variations

  • Add a category once real traffic shows a cluster the existing ones don’t cover: a new entry in the routing table, not a new technique.
  • Move to function calling once the queue depends on a lookup the label can’t settle, and to a single agent once that lookup takes a different number of steps each time.
  • Retune the gate from what actually went wrong. Human approval’s own advice is to set the threshold from the answers that turned out wrong before, rather than from a guess about which ones will.

Design choices

Why this level, and when to use another approach

Three techniques compose this recipe: routing reads the message and picks one of a handful of categories your code already wrote a handler for; structured output turns what the message says into a fixed-shape ticket record instead of prose a downstream system would have to parse; human approval pauses before filing anything the classifier wasn’t confident about, or anything naming a cost.

Level 3 is enough because both halves of the job are known in advance: the categories are fixed (billing, support, general, and a fallback for anything that fits none of them), and so is what happens once a message lands in one. The routing page draws exactly this line: a rule is enough when the categories are easy to tell apart from the wording, and a classifier earns its keep once the wording gets too varied for a rule to catch reliably. The function calling page states the same boundary from the other side: try routing first whenever the input’s surface form already tells you which single action applies. That is this job. The label picks the queue, so a tool would only let the model re-make a choice your code has already made, at the cost of the extra call function calling’s own figures show on the tool branch.

The human-in-the-loop step mirrors its own example’s two rules almost exactly: no confident category (the routing table’s own fallback) or a named dollar amount (the same high_cost check that page’s example uses) pauses the run. It stays a level-3 decision for the reason that page gives directly: your code decides when to pause, against a fixed rule, and the model is never asked whether a person should look. Which rule you pick is the whole of it: too loose and the gate waves through what most needed a look, too tight and approving turns into a reflex.

Moving higher only pays off once the right queue depends on something the message’s own text can’t settle: an account lookup, or a search that might take one step or several. Function calling is the first of those, a single agent the next. Neither is needed to sort a message that already says everything the router needs to know.

Composition

Techniques this recipe uses

The highest level it needs is level 3.

Routing

Sourced

Sorting inputs and sending each one to the right prompt.

Structured output

Sourced

Getting answers in a fixed format such as JSON.

Human approval

Sourced

Pausing for a person to approve or correct.

Same shape, other jobs

Sort incoming items and send each where it belongs

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • Tenant, customer or patient messages by urgency
  • Support tickets by product area
  • Failed units by likely cause: fixture, lot, handling or design
  • Bug reports by component and severity
  • Monitoring alerts by who should be paged
  • Incoming leads by fit

Last reviewed 09/18/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page