Recipe

Support desk

Routes an incoming ticket, searches the documentation for an answer, calls a tool when an action is needed, and hands off to a person when it is unsure.

SourcedNeeds level 4

A hardware shop runs a support desk: customers write in about a part that failed, a warranty question, or something that doesn’t work the way the manual says. Someone reads each ticket, checks the documentation and, for the tickets that need it, issues a replacement rather than answering in words.

Support desk composes that into one pipeline: sort the ticket, search the documentation, let the model reach for a tool when the ticket needs an action rather than an answer, and pause before anything costly goes out.

Example run

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Support desk, assembled

Route the ticket, search the docs, let the model reach for a tool if it needs one, and pause before anything costly goes out.

Level 4 · Tool use
Ticket arrivesTicket arrivesMODELClassify the ticketClassify the ticketRoute to the technical queueRoute to thetechnical queueSearch the docsSearch the docsMODELAnswer, or call a toolAnswer, orcall a toolTOOLissue_replacement(part)issue_replacement(part)MODELDraft an answerDraft an answerCheck confidenceCheck confidencePERSONPerson decidesPerson decidesResumeResumeTicket answeredTicket answered
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 10Your code chose

The ticket arrives

"The drain pump on my DW-300 (HLV-2205) failed under
warranty, still rattling after I replaced it once already."
0 tokens · 0 ms

Walkthrough

The four steps compose the runnable code already on the routing, RAG, function calling and human approval pages, each written against the site’s shared document set rather than a literal support desk. Composing them means keeping the same functions and swapping in this job’s own queues and tool: billing / technical / general instead of lookup / numeric / unclear, and issue_replacement(part, reason) in place of lookup_part(part_number). The tool is the only genuinely new code; no version of it exists in the repository yet.

The run shown is a repeat-failure ticket. It classifies technical, and the search step retrieves the warranty policy and the part’s own listing. The model reads both and decides on its own that a repeat failure under warranty calls for a real replacement, and calls the tool with the part number and its reasoning. Your code runs the tool and the model drafts the reply from the result: the same decided_by: "model" step, and the same “your code always runs whatever it calls” boundary, that function calling’s own page describes. Because a tool ran, the gate pauses on tool_used before the reply goes out, and a person checks the ticket, the sources and the tool call together before it ships.

Be plain about two things. The ticket text and the retrieved documentation both go to the model, so a desk handling anything customers would not want copied elsewhere has a hosting decision to make before an accuracy one. And log the tools offered, the call and its arguments, the retrieved sections, the gate’s reason and the reviewer’s decision together, or a replacement that should never have shipped cannot be traced to the ticket, the sources or the argument the model chose.

What to measure

Routing, RAG and function calling are each scored on their own by the site’s shared 60-question set (see docs/EVALS.md); this job’s version of each (is the ticket sorted the way a person would, does the citation hit rate on retrieved sources hold up, does model_decided_steps come back as exactly one per ticket, never zero and never more) reuses those numbers, on a labeled set of real or synthetic tickets instead of document questions. Watch one thing routing and RAG alone can’t tell you: whether the tool actually gets called on the tickets that need an action, and only those: a model that reaches for it on every ticket, or never, has learned the wrong lesson from its description. Human approval isn’t scored by the shared runner (a paused ticket has no answer to grade) and needs the three numbers its own page describes instead: pause rate, whether the right tickets pause, and accuracy after a person’s decision. No result file exists for any of the four, so this recipe claims no score.

Variations

  • Reach for MCP once more than one application needs these same actions, or the tools should come from a server the support team doesn’t maintain.
  • Move to knowledge graphs if tickets start needing facts joined across documents: a part’s compatible models and each model’s own warranty class, answered together rather than by one lucky search.
  • Move to agentic RAG if a single search often misses what a second, better-aimed one would find, before the tool decision happens.
  • Add a second tool (a refund, an escalation to a technician) once real tickets show a second action worth automating; the routing table and the gate carry over unchanged.

Design choices

Why this level, and when to use another approach

Four techniques compose this recipe. Routing sorts the ticket into a fixed queue. RAG searches the shop’s documentation and drafts an answer from what it finds, cited. Function calling is the one place this recipe climbs past a fixed workflow: the model is offered a tool (issue a replacement part) and decides for itself whether this ticket needs it, rather than your code deciding in advance from the queue label alone. Human approval pauses before a reply that used the tool goes out, since a free replacement has a real cost if the model called it on a ticket that didn’t actually warrant one.

Level 4 is the right stop, not level 3, because the thing being decided here is what function calling’s own page says the level is for: whether this ticket needs a replacement shipped rather than an answer written is a judgment about open-ended text that a fixed rule can’t make reliably. Routing can pick the queue; it can’t decide whether to act. That one model decision is all this recipe adds over a plain sort-and-answer workflow, and it costs accordingly: function calling’s own figures are two model calls on a ticket that gets a tool call, against RAG’s one.

It is not worth climbing to a single agent unless the tool decision itself needs more than one step (the function calling page states that upgrade condition directly): move up once the next action depends on what the last one returned, so the model has to choose again and decide when to stop. This desk’s tool call is a single bounded action with a known result; nothing asks the model to check the outcome and decide again. A ticket that needed the model to check an order’s shipping status first and act differently on what that showed is the case a single agent is for, at several times the cost.

Composition

Techniques this recipe uses

The highest level it needs is level 4.

Routing

Sourced

Sorting inputs and sending each one to the right prompt.

Retrieval-augmented generation (RAG)

Measured

Searching your documents and giving the results to the model.

Function calling

Sourced

Letting the model call functions that you define.

Human approval

Sourced

Pausing for a person to approve or correct.

Same shape, other jobs

Answer people in conversation, looking things up and taking small actions

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • Customer support
  • An internal IT or HR help desk
  • Booking shared lab equipment and checking its calibration status
  • Order and delivery status
  • A parts-availability assistant for a purchasing team

Last reviewed 09/18/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page