# Sort an inbox

_Recipe · needs level 3_

Sorts mail into fixed categories and produces structured output. A person approves anything that gets sent. The categories are known in advance, so an agent is not needed.


A small company's shared inbox gets everything: billing questions, product-support requests,
general questions, and the occasional message that fits none of those. Someone has to read each
one, decide where it goes, and get the details into the system that handles that category,
instead of leaving it as an email somebody must open and reread field by field. Nobody wants a
refund request quietly filed into a queue with no person ever seeing the dollar figure in it.

Inbox triage does that in three fixed steps: classify the message into one of a known, small set
of categories, turn it into a structured record, and file it to the matching queue. Anything the
classifier isn't confident about, or that names a dollar amount, waits for a person first.

## Example run

_The web page for this technique includes an interactive step-through of Level 3 · Sort an inbox. The same steps are described in the sections below._

## Walkthrough

The three steps compose the runnable code already on each technique's own page, unmodified in
shape: `examples/routing/run.py`'s classify-then-dispatch pattern, `examples/structured_output/
run.py`'s ask-validate-retry-once pattern, and `examples/human_in_the_loop/run.py`'s
check-thresholds-and-resume pattern. All three are written against the site's shared document set,
not a literal inbox, so this exact version is not in the repository. Composing them means keeping
the same functions and swapping in this job's own labels (`billing`, `support`, `general`,
`unclear` instead of `lookup`, `numeric`, `unclear`) and its own schema, a ticket record instead
of a warranty record. The parsing, the routing table and the one-retry contract carry over
unchanged; none of them inspect what the labels are called.

The run above follows one message: a refund request naming a part and a dollar amount.
It classifies as `billing`, and the handler extracts a ticket record (category, summary, amount,
priority), the same shape structured output's own example validates before accepting. The gate
trips on `high_cost`, since the record names a dollar figure, and the run pauses rather than
filing it. A person sees the message and the record, approves it, and the code files the ticket.
A message that classifies `unclear` pauses by the same code path for the opposite reason: the
fallback found nothing confident enough to hand a handler at all.

Two things to settle before this runs on real mail. Every message body goes to whichever model
classifies it, so where that model runs is a decision about customer data and not only about
accuracy. And log the raw classification text, the parsed label, the record, the gate's reason and
the person's decision together: an approval with no record of what the reviewer was shown cannot
be audited afterwards.

## What to measure

Build a labeled set for this job: fifty to a hundred real or synthetic messages, hand-labeled with
the category and the ticket fields a person would write down. Three numbers, each testing a
different piece: routing accuracy (does the label match the human one); extraction accuracy field
by field, with the valid-JSON rate and retry count tracked separately; and the pause rate split by
reason, checked against a sample of messages that did *not* pause, to see whether any should have.
No result file exists for routing, structured output or human approval yet (see
`docs/EVALS.md`), so this recipe claims no score.

## Variations

- Add a category once real traffic shows a cluster the existing ones don't cover: a new entry
  in the routing table, not a new technique.
- Move to [function calling](/gradient_ascent/techniques/function-calling/) once the queue
  depends on a lookup the label can't settle, and to [a
  single agent](/gradient_ascent/techniques/single-agent/) once that lookup takes a different number of steps each time.
- Retune the gate from what actually went wrong. [Human approval](/gradient_ascent/techniques/human-in-the-loop/)'s own advice is to set the threshold
  from the answers that turned out wrong before, rather than from a guess about which ones will.

## Design choices

### Why this level, and when to use another approach

Three techniques compose this recipe: [routing](/gradient_ascent/techniques/routing/) reads the
message and picks one of a handful of categories your code already wrote a handler for;
[structured output](/gradient_ascent/techniques/structured-output/) turns what the message says
into a fixed-shape ticket record instead of prose a downstream system would have to parse;
[human approval](/gradient_ascent/techniques/human-in-the-loop/) pauses before filing anything
the classifier wasn't confident about, or anything naming a cost.

Level 3 is enough because both halves of the job are known in advance: the categories are fixed
(billing, support, general, and a fallback for anything that fits none of them), and so is what
happens once a message lands in one. The routing page draws exactly this line: a rule is enough
when the categories are easy to tell apart from the wording, and a classifier earns its keep once
the wording gets too varied for a rule to catch reliably. The function calling page states the
same boundary from the other side: try routing first whenever the input's surface form already
tells you which single action applies. That is this job. The label picks the queue, so a tool
would only let the model re-make a choice your code has already made, at the cost of the extra
call function calling's own figures show on the tool branch.

The human-in-the-loop step mirrors its own example's two rules almost exactly: no confident
category (the routing table's own fallback) or a named dollar amount (the same `high_cost` check
that page's example uses) pauses the run. It stays a level-3 decision for the reason that page
gives directly: your code decides when to pause, against a fixed rule, and the model is never
asked whether a person should look. Which rule you pick is the whole of it: too loose and the gate
waves through what most needed a look, too tight and approving turns into a reflex.

Moving higher only pays off once the right queue depends on something the message's own text
can't settle: an account lookup, or a search that might take one step or several. [Function calling](/gradient_ascent/techniques/function-calling/) is the first of those, [a single agent](/gradient_ascent/techniques/single-agent/) the next. Neither is needed to sort a
message that already says everything the router needs to know.



Last reviewed 2026-09-18.
