The example classifies a question as lookup, numeric or unclear with one small model call,
then dispatches to one of three fixed handlers. Only the lookup handler calls the model again;
the numeric handler answers from an exact part-number match in the parts list with no model call
at all, and the fallback answers nothing on purpose rather than guess. A question that reaches
the cheapest handler that can actually answer it costs less than one that reaches the most
capable one by default: the point Anthropic makes about routing to a smaller model for easy
questions[1] generalizes to routing to no model at all when a plain lookup will do.
_parse_label is the whole boundary between what the model decided and what the code decided:
it takes the model’s raw text, keeps only the first word, and returns it if and only if it is one
of the three known labels: anything else, including a hedge like “probably lookup,” becomes
unclear. The routing table itself, ROUTES = {"lookup": ..., "numeric": ..., "unclear": ...},
is a plain dict the code wrote before the first question ever arrived. The model can steer which
value comes out of _parse_label; it cannot add a fourth key to ROUTES.
The fallback route matters as much as the working ones. _numeric_route calls the fallback
itself when the label says “numeric” but no part number pattern actually appears in the
question. The classifier can be confident about the wrong thing, and the honest response is to
defer rather than to force an answer out of a handler that has nothing to work with.
examples/routing/run.py · lines 23–96
LEVEL = 3
LOOKUP_K = 3
PART_RE = re.compile(r"HLV-\d{4}")
LABELS = ("lookup", "numeric", "unclear")
CLASSIFY_SYSTEM = (
"Classify the question as exactly one word: 'lookup' if it asks about a fact described in a "
"Halvorsen document, 'numeric' if it asks for one part's price or part number, or 'unclear' "
"if it is neither, or you are not confident. Reply with exactly one of those three words."
)
LOOKUP_SYSTEM = (
"You answer questions about Halvorsen appliances using only the numbered sources below. End "
"your answer with a line starting 'Sources:' listing the citations, like 'dw300-manual#3'."
)
def _parse_label(text: str) -> str:
first_word = text.strip().split()[0].lower().strip(".,:;\"'") if text.strip() else ""
return first_word if first_word in LABELS else "unclear"
def _lookup_route(question: str, sections: dict[str, Section], model: Model, tracer: Tracer) -> Answer:
sources = [s for s, score in bm25_search(sections, question, k=LOOKUP_K) if score > 0]
blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
completion = model.complete(
[Message(role="system", content=LOOKUP_SYSTEM), Message(role="user", content=f"Sources:\n\n{blocks}\n\nQuestion: {question}")],
max_tokens=400,
)
tracer.record(
kind="model", decided_by="code", title="Answer with the lookup prompt", detail=completion.text[:200],
tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
)
return Answer.from_text(completion.text, retrieved_sources=[s.cite for s in sources])
def _numeric_route(question: str, sections: dict[str, Section], model: Model, tracer: Tracer) -> Answer:
match = PART_RE.search(question.upper())
if not match:
tracer.record(kind="code", decided_by="code", title="Numeric route found no part number", detail="falling back to ask a person")
return _person_route(question, sections, model, tracer)
line = lookup_part(match.group(0))
tracer.record(kind="code", decided_by="code", title="Answer with the numeric route", detail=line or f"{match.group(0)} not found")
if not line:
return Answer(text=f"{match.group(0)} is not in the parts list.", citations=[])
cite = next((c for c, s in sections.items() if c.startswith("parts-list") and match.group(0) in s.text), None)
return Answer(text=line, citations=[cite] if cite else [], retrieved_sources=[cite] if cite else [])
def _person_route(question: str, sections: dict[str, Section], model: Model, tracer: Tracer) -> Answer:
del question, sections, model # the fallback answers nothing; it defers, on purpose
tracer.record(kind="code", decided_by="code", title="Route to a person", detail="no automatic route was confident enough")
return Answer(text="This needs a person to check; no automatic route here was confident enough to answer it.", citations=[])
ROUTES = {"lookup": _lookup_route, "numeric": _numeric_route, "unclear": _person_route}
def run(
question: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer:
del embedder # routing retrieves by keyword inside the lookup route, not by vector
sections = load_sections(corpus_dir)
classify = model.complete([Message(role="system", content=CLASSIFY_SYSTEM), Message(role="user", content=question)], max_tokens=5)
tracer.record(
kind="model", decided_by="code", title="Classify the question", detail=classify.text.strip(),
tokens_in=classify.tokens_in, tokens_out=classify.tokens_out, ms=classify.ms,
)
label = _parse_label(classify.text)
tracer.record(kind="code", decided_by="code", title="Route on the label", detail=f"label={label!r} -> {label} route")
return ROUTES[label](question, sections, model, tracer)
Run it yourself:
examples/routing/README.md · lines 15–15
python -m examples.routing --model stub:scripted
The classify step does not have to be a model call. Semantic Router, a library built around this
one pattern, compares the question’s embedding against a few example utterances per route and
picks the closest, which its own README describes as making the decision in semantic vector space
rather than waiting for a model to generate it[3]. That is the same boundary drawn in a
cheaper place: a table of routes your code wrote, and a classifier that can only choose among
them.
A router is worth measuring on its own, separately from whether the final answer was right. Feed
it a small set of questions you have hand-labeled with the intended route, and score the
classify step alone: what share got the label a person would have picked. A router that is 95%
accurate but only used 60% of the time (because most traffic quietly falls to “unclear”) is a
different problem than one that is used 95% of the time but wrong on a fifth of what it routes,
and a single end-to-end accuracy number cannot tell those apart.