# Plain-language maintenance log

_Recipe · needs level 4_

Turns a plain-language description of work done into a structured log entry, saved with a tool call and linked to the equipment it concerns through a small knowledge graph.


A small facility's technicians write what they did in plain language: "replaced the belt on pump
3, it was fraying", rather than filling out a rigid form. Someone still needs that turned into a
real log entry: which piece of equipment, what was done, when, filed against that equipment's own
history rather than as a flat note nobody can search. And every so often, a note buried in
otherwise routine language is actually describing something that needs a person's attention now,
not at the next scheduled review.

The pipeline is three steps: extract a fixed record from the note, find the equipment it concerns
in a small graph of what the facility has, and let the model decide whether this is routine or
needs flagging.

## Example run

_The web page for this technique includes an interactive step-through of Level 4 · Plain-language maintenance log. The same steps are described in the sections below._

## Walkthrough

The three steps compose the runnable code already on the structured output, knowledge graphs and
function calling pages, unmodified in shape, each written against the site's own shared examples
(structured output reads a warranty record, knowledge graphs walks from a part to its warranty
class) rather than a literal maintenance note. Composing them means keeping the same functions
and swapping in this job's own schema and its own two-hop walk (equipment to line, line to
last-serviced date), and offering `save_log_entry` and `flag_for_review` in place of
`lookup_part`. Nothing here ships that exact pair of tools yet; what's illustrated above is the
shape those three pages' code already is, run for this job.

The run above shows a note that reads as routine until its last clause. Extraction pulls a clean
record (equipment, action, a condition field carrying the technician's own words) and the graph
walk finds pump 3's line and that it was serviced two months ago, nothing unusual on its own. The
model reads the condition field's "starting to smell hot" alongside a worn belt and calls
`flag_for_review` instead of `save_log_entry`, the one `decided_by: "model"` step in the whole
run; a note that only said "replaced the belt, routine wear" would have taken the other tool
instead, with nothing else in the pipeline changing. Either way your code links the entry into the
equipment's history; what the flag adds is a review task on top of it. Log both tools offered,
which one was called, and the graph path behind the record, so a note that should have been
flagged and wasn't can be found afterwards.

## What to measure

Knowledge graphs and function calling are each scored on their own by the site's shared
60-question set (see `docs/EVALS.md`); structured output is not. None of the three answers this
job's actual task, so build a labeled set of real or synthetic notes instead: the record a person
would extract, the equipment it should link to, and whether a facilities lead would flag it. Watch
whether the escalation tool gets called on the notes that actually need it, and only those,
checked against a sample of routine entries to confirm none should have been flagged instead. Add
one measure specific to the graph: how often the named equipment fails to resolve to anything in
it, which is a data problem, not a model one. No result file exists for any of the three on this
task yet, so this recipe cannot claim a score for any of it.

## Variations

- Add [human approval](/gradient_ascent/techniques/human-in-the-loop/) before saving an entry
  whose equipment id the graph doesn't recognize: a missing link and a safety concern are
  different problems and shouldn't share one gate.
- Move to [a single agent](/gradient_ascent/techniques/single-agent/) if escalating ever needs
  more than one lookup: checking recent notes on similar equipment before judging this one.
- Route by facility area first with [routing](/gradient_ascent/techniques/routing/) if the same
  pipeline serves more than one site, each with its own equipment graph.
- Rebuild the graph on a schedule, not on every note; knowledge graphs' own page is direct about
  the cost: extraction is expensive, and paid once per equipment change, not per entry.

## Design choices

### Why this level, and when to use another approach

Three techniques compose this recipe. [Structured
output](/gradient_ascent/techniques/structured-output/) turns the note into a fixed record (equipment, action, a condition field, a date)
validated before anything downstream touches it. [Knowledge graphs](/gradient_ascent/techniques/knowledge-graphs/) is what makes "pump 3" mean something:
a small graph walk from the named equipment to what line it's on and when it was last serviced, so
the entry lands linked to a real history rather than as an isolated row. [Function calling](/gradient_ascent/techniques/function-calling/) is the one place this recipe climbs
past a fixed pipeline: the model is offered two tools (log it routinely, or flag it for review)
and decides which one this note actually calls for.

That decision is why level 4, not level 2, is the honest floor here. Structured output and
knowledge graphs alone would produce a well-formed, well-linked record every time, but they'd file
"smelled hot" with the same routine handling as "replaced the belt": neither technique reads for
risk, only for fields and connections. Function calling's own page draws the boundary this recipe
needs: the model is offered a schema-shaped action and decides on its own whether to reach for it.
Structured output's page draws the same contrast from its own side: at level 4 the model
additionally decides *whether* to use a shape at all, not just what goes in it. A fixed keyword
list could
catch a few obvious cases, but "starting to smell hot" is exactly the open-ended phrasing such a
list can't be written to catch reliably in advance.

It isn't worth climbing past that single decision. The model acts once, on one note, with
everything it needs already in the extracted record and the graph walk; nothing here asks it to
check a result and decide again. A facility that wanted the model to weigh a piece of equipment's
full service history against similar recent notes elsewhere before deciding (several lookups,
each depending on the last) is the case a single agent is for, at real added cost for a decision
this recipe currently makes in one call.



Last reviewed 2026-09-18.
