# Routing

_Level 03 · Workflows · sourced_

Sorting inputs and sending each one to the right prompt.


## Try this in a recipe
- [Build a weekly update without invented progress](/gradient_ascent/recipes/weekly-status-report.md): Extract evidence into a checked table, then draft an update from that table in a fixed two-call workflow.

## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow an incoming request into one of several paths. Inspect the evidence for that choice, especially when the request fits more than one category or none clearly.

**Assumptions:** The available destinations must have meaningful responsibilities. A forced label can hide ambiguity or a missing route.

**Design choices:** Use explicit rules for obvious cases and model classification for language variation when useful. Allow clarification, multiple labels, or an unresolved queue if the task needs them.

**Request:** Route this support request to the appropriate queue.

**Starting evidence:** Message: I was charged twice and my device will not start. Queues: billing, technical support, human triage.

**Action and control:** Identify two intents and use the mixed-intent route instead of discarding an issue.

**Stage records (authored, not executed):**

### Input record

Message: I was charged twice and my device will not start. Queues: billing, technical support, human triage.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use explicit rules for obvious cases and model classification for language variation when useful. Allow clarification, multiple labels, or an unresolved queue if the task needs them.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Identify two intents and use the mixed-intent route instead of discarding an issue.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Route: human triage or linked billing and technical tickets according to policy. Both concerns retained.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

A routing decision, confidence limitation, mixed-intent case, and a confusion matrix on labeled sample messages.

If the result falls short:
When routing is uncertain or wrong, preserve the original request and offer a correction path. Track costly misroutes rather than accuracy alone.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to personal inboxes, support, engineering triage, or choosing tools. Your categories and escalation threshold should reflect who handles the work next.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Route: human triage or linked billing and technical tickets according to policy. Both concerns retained.

**Change something — Remove the mixed-intent route:** Escalate the ambiguous case; billing alone silently drops the technical issue.

**Decision:** Should one label silently erase the second concern?

**Answer:** No; preserve it or escalate.

**Why:** Mixed intent and low confidence require explicit handling; a route choice does not resolve the underlying issue.

**Review criteria:** A routing decision, confidence limitation, mixed-intent case, and a confusion matrix on labeled sample messages.

**Recovery:** When routing is uncertain or wrong, preserve the original request and offer a correction path. Track costly misroutes rather than accuracy alone.

**Adapt it:** Apply this to personal inboxes, support, engineering triage, or choosing tools. Your categories and escalation threshold should reflect who handles the work next.

Routing looks at an input, decides which of several fixed kinds it is, and sends it down the
path built for that kind. Anthropic's description is "Routing classifies an input and directs it
to a specialized followup task", which it says allows separation of concerns and "building more
specialized prompts"[1]. Its second example spends the idea on cost instead: easy or
common questions to smaller, cost-efficient models, hard or unusual ones to more capable
ones[1].

Routing sits at level 3, workflows. The classifying step is a code-owned decision even though a
model produces the label: the model answers a narrow question with one word, and your code looks
that word up in a table of handlers it wrote in advance. The line to level 4 falls there. At
level 4 the model is offered a tool and its output invokes one directly; here the label is only a
value your code branches on.

A model built for exactly this step appeared in September 2026. TypeSafe AI's Jev returns a typed
choice instead of text, with what TypeSafe calls "calibrated probabilities and confidence
scores"[4]. A probability is worth more to a router than a bare label. Jev is in early
access, and that is TypeSafe's claim, not a measurement here.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 3 · Routing. The same steps are described in the sections below._

## Practical guidance

Build a router in an automation tool with a visual canvas: a trigger, then a branch step that
sends different kinds of item down different paths. Zapier's Paths feature is a direct product
version of this: "if 'A' happens in your first app, then do this, but if 'B' happens, do
something else"[2], with a rule for each path controlling what is allowed to reach it.

Start with a plain rule for the categories a keyword or a field value already tells apart, such
as sending anything whose subject line contains "refund" to billing. Add a model-based branch
only for the categories wording alone cannot reliably sort: point it at a step that reads the
whole item and returns one of a fixed list of labels you already named on the canvas, such as
billing, technical or general, and wire each label to its own path. A model-based branch is more
forgiving of wording you did not anticipate than a rule, and more expensive and less predictable,
since two similar items can occasionally get different labels.

Give every branch a real destination, including the one for "none of these." If unclear or other
quietly becomes the largest path, the router is not routing, it is mostly declining, and whoever
or whatever sits on that path needs to actually handle it rather than let items pile up unseen.
That is where [a person approving](/gradient_ascent/techniques/human-in-the-loop/) earns its
place.

Open a branch step and read exactly what decided it, the same way you would check a step in a
chain. A rule shows its condition in plain text on the canvas. A model-based branch shows you the
label it returned; if the tool also shows a confidence score, send anything below a threshold you
pick to a person instead of down whichever path the label happened to name.

Once it is running, periodically pull a sample of whatever landed in the fallback path and check
whether that share is growing under real traffic. A fallback that starts small and quietly
becomes the biggest path is the router breaking down, not real traffic getting harder.

## Implementation details

The example classifies a question as `lookup`, `numeric` or `unclear` with one small model call,
then dispatches to one of three fixed handlers. Only the lookup handler calls the model again;
the numeric handler answers from an exact part-number match in the parts list with no model call
at all, and the fallback answers nothing on purpose rather than guess. A question that reaches
the cheapest handler that can actually answer it costs less than one that reaches the most
capable one by default: the point Anthropic makes about routing to a smaller model for easy
questions[1] generalizes to routing to no model at all when a plain lookup will do.

`_parse_label` is the whole boundary between what the model decided and what the code decided:
it takes the model's raw text, keeps only the first word, and returns it if and only if it is one
of the three known labels: anything else, including a hedge like "probably lookup," becomes
`unclear`. The routing table itself, `ROUTES = {"lookup": ..., "numeric": ..., "unclear": ...}`,
is a plain dict the code wrote before the first question ever arrived. The model can steer which
value comes out of `_parse_label`; it cannot add a fourth key to `ROUTES`.

The fallback route matters as much as the working ones. `_numeric_route` calls the fallback
itself when the label says "numeric" but no part number pattern actually appears in the
question. The classifier can be confident about the wrong thing, and the honest response is to
defer rather than to force an answer out of a handler that has nothing to work with.

`examples/routing/run.py` (lines 23-96)

```python
LEVEL = 3
LOOKUP_K = 3
PART_RE = re.compile(r"HLV-\d{4}")
LABELS = ("lookup", "numeric", "unclear")
CLASSIFY_SYSTEM = (
    "Classify the question as exactly one word: 'lookup' if it asks about a fact described in a "
    "Halvorsen document, 'numeric' if it asks for one part's price or part number, or 'unclear' "
    "if it is neither, or you are not confident. Reply with exactly one of those three words."
)
LOOKUP_SYSTEM = (
    "You answer questions about Halvorsen appliances using only the numbered sources below. End "
    "your answer with a line starting 'Sources:' listing the citations, like 'dw300-manual#3'."
)

def _parse_label(text: str) -> str:
    first_word = text.strip().split()[0].lower().strip(".,:;\"'") if text.strip() else ""
    return first_word if first_word in LABELS else "unclear"

def _lookup_route(question: str, sections: dict[str, Section], model: Model, tracer: Tracer) -> Answer:
    sources = [s for s, score in bm25_search(sections, question, k=LOOKUP_K) if score > 0]
    blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
    completion = model.complete(
        [Message(role="system", content=LOOKUP_SYSTEM), Message(role="user", content=f"Sources:\n\n{blocks}\n\nQuestion: {question}")],
        max_tokens=400,
    )
    tracer.record(
        kind="model", decided_by="code", title="Answer with the lookup prompt", detail=completion.text[:200],
        tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )
    return Answer.from_text(completion.text, retrieved_sources=[s.cite for s in sources])

def _numeric_route(question: str, sections: dict[str, Section], model: Model, tracer: Tracer) -> Answer:
    match = PART_RE.search(question.upper())
    if not match:
        tracer.record(kind="code", decided_by="code", title="Numeric route found no part number", detail="falling back to ask a person")
        return _person_route(question, sections, model, tracer)
    line = lookup_part(match.group(0))
    tracer.record(kind="code", decided_by="code", title="Answer with the numeric route", detail=line or f"{match.group(0)} not found")
    if not line:
        return Answer(text=f"{match.group(0)} is not in the parts list.", citations=[])
    cite = next((c for c, s in sections.items() if c.startswith("parts-list") and match.group(0) in s.text), None)
    return Answer(text=line, citations=[cite] if cite else [], retrieved_sources=[cite] if cite else [])

def _person_route(question: str, sections: dict[str, Section], model: Model, tracer: Tracer) -> Answer:
    del question, sections, model  # the fallback answers nothing; it defers, on purpose
    tracer.record(kind="code", decided_by="code", title="Route to a person", detail="no automatic route was confident enough")
    return Answer(text="This needs a person to check; no automatic route here was confident enough to answer it.", citations=[])

ROUTES = {"lookup": _lookup_route, "numeric": _numeric_route, "unclear": _person_route}

def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer:
    del embedder  # routing retrieves by keyword inside the lookup route, not by vector
    sections = load_sections(corpus_dir)
    classify = model.complete([Message(role="system", content=CLASSIFY_SYSTEM), Message(role="user", content=question)], max_tokens=5)
    tracer.record(
        kind="model", decided_by="code", title="Classify the question", detail=classify.text.strip(),
        tokens_in=classify.tokens_in, tokens_out=classify.tokens_out, ms=classify.ms,
    )
    label = _parse_label(classify.text)
    tracer.record(kind="code", decided_by="code", title="Route on the label", detail=f"label={label!r} -> {label} route")
    return ROUTES[label](question, sections, model, tracer)
```

Run it yourself:

`examples/routing/README.md` (lines 15-15)

```text
python -m examples.routing --model stub:scripted
```

The classify step does not have to be a model call. Semantic Router, a library built around this
one pattern, compares the question's embedding against a few example utterances per route and
picks the closest, which its own README describes as making the decision in semantic vector space
rather than waiting for a model to generate it[3]. That is the same boundary drawn in a
cheaper place: a table of routes your code wrote, and a classifier that can only choose among
them.

A router is worth measuring on its own, separately from whether the final answer was right. Feed
it a small set of questions you have hand-labeled with the *intended* route, and score the
classify step alone: what share got the label a person would have picked. A router that is 95%
accurate but only used 60% of the time (because most traffic quietly falls to "unclear") is a
different problem than one that is used 95% of the time but wrong on a fifth of what it routes,
and a single end-to-end accuracy number cannot tell those apart.

## When you do not need this

Try a plain rule (a keyword, a regular expression, a dropdown the user picks from) first if the
categories are few and easy to tell apart from the surface form of the input: a rule is free to
run, free to test, and never drifts between two similar inputs the way a classifier can. That is
[level 0, no model at all](/gradient_ascent/techniques/order-zero/).

A rule and a classifier are not a choice of one. Where one category is dangerous to miss, run
both and escalate if either one fires. The rule catches the plain cases for nothing ("gas",
"smoke", "flooding") and keeps catching them on the day the classifier gets one wrong. The
classifier catches the tenant who writes "something smells odd by the stove". Each covers the
other's blind spot, and the cost is a few lines of code.

Move up to routing once the categories are real but the wording that signals each one is too
varied to write as a rule, and a wrong route is cheap enough to tolerate at the rate a classifier
gets it wrong.

## Failure modes

### Confident misroute

- **How to notice it:** The classifier names a label with no hedge, the handler runs, and the answer is fluent and wrong, because the input actually needed a different route than the one it confidently got.
- **How to test for it:** Score the classify step alone against a hand-labeled set of questions and their intended routes, separately from whether the final answer was correct, so a wrong route and a wrong answer from a right route are not the same number.

### The fallback route is missing or too weak

- **How to notice it:** "unclear" or an unrecognized label reaches a handler that guesses anyway instead of declining, because the fallback path was never given as much attention as the main ones.
- **How to test for it:** Send it questions built to be genuinely ambiguous and confirm the fallback route actually defers, rather than picking one of the other handlers by default.

### Category drift

- **How to notice it:** The share of questions landing in each category shifts over time (a new kind of question starts arriving that fits none of the categories well), and the router keeps forcing it into the closest existing one.
- **How to test for it:** Track the label distribution over time, not just per-run accuracy; a category whose share moves a lot without a matching shift in the real input mix is worth a manual sample.

### A route that is cheaper but does not actually answer

- **How to notice it:** The cheap, model-free route (a lookup table, a fixed rule) is chosen because the label matched, but the specific case is one that route cannot really handle, so it returns a technically-on-topic but wrong or incomplete answer.
- **How to test for it:** Check the numeric route specifically against questions naming a part number pattern that is not actually in the parts list, and confirm it reports 'not found' rather than inventing a price.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, lookup route:** 2
- **Model calls, numeric or unclear route:** 1
- **Classify tokens in:** ~64
- **Classify tokens out:** 1

**Compared with RAG (level 2), every question.** RAG spends one full retrieval-and-answer call on every question regardless of kind. Routing spends a small classification call on every question, but only the questions routed to the lookup handler pay for a second, larger call.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

Two numbers matter here, not one. End-to-end accuracy on the same 60-question set as every other
technique is the first; routing accuracy (whether the classify step's label matches the kind a
person would assign the question) is the second, and the site scores them separately so a wrong
route and a right route with a wrong answer are not confused with each other.

The `numeric` kind maps directly onto the example's part-number handler; `unanswerable`
questions are the clearest test of the fallback, since the right behavior is to decline rather
than force an answer through whichever handler the classifier happened to name. No result file
exists yet (see `docs/EVALS.md`). Run `python scripts/eval_run.py --example routing --model
<spec> --dry` to project the cost of a real run before spending anything on one.

## Run it

**What to monitor.** The label distribution over time, and accuracy of the classify step against a small hand-labeled sample, tracked separately from end-to-end answer accuracy. A route whose share of traffic changes sharply, with no matching change in the real input mix, is worth a manual look.

**Cost at volume.** Every question pays for one small classification call; only the questions routed to a model-calling handler pay for a second, larger one. Cost tracks the mix of routes actual traffic takes, not a fixed per-question number the way a single-path technique's cost does.

**How it fails in production.** A new kind of question starts arriving that fits none of the categories, and the classifier keeps forcing it into the closest existing label instead of the fallback, because nothing told it that kind did not exist yet when it was built.

**What to log.** The question, the raw classification text before parsing, the parsed label, and which handler actually ran, so a wrong answer can be traced to a misclassification, a parsing bug, or a handler that ran correctly on the wrong input.

## Try it

1. **Use it.** Find a form or a support inbox that already sorts incoming items into a few fixed categories. Write down what happens to something that fits none of them well. Is there a real fallback, or does it get forced into the closest category?
2. **Build it.** Run python -m examples.routing --model stub:scripted from the repo root. The classifier returns lookup, the code routes on that one word, and the lookup route answers with its citation. Run it again with --model stub: the echo is not one of the labels, so every question lands in the unclear fallback and gets the same answer, which is what a classifier that never returns a label looks like from the outside.
3. **Either lane.** Write five questions on purpose to land in the "unclear" fallback. How easy was that? A route that is too easy to fall into is a sign the other categories are drawn too narrowly.


## Sources

1. [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic, 2024-12-19 (accessed 2026-09-19)
2. [Paths](https://zapier.com/features/paths) — Zapier (accessed 2026-09-19)
3. [semantic-router](https://github.com/aurelio-labs/semantic-router) — Aurelio Labs (GitHub README) (accessed 2026-09-19)
4. [Introducing System One Models & Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) — TypeSafe AI, 2026-09-15 (accessed 2026-09-19)


Last reviewed 2026-09-19.
