Recipe

Plain-language maintenance log

Turns a plain-language description of work done into a structured log entry, saved with a tool call and linked to the equipment it concerns through a small knowledge graph.

SourcedNeeds level 4

A small facility’s technicians write what they did in plain language: “replaced the belt on pump 3, it was fraying”, rather than filling out a rigid form. Someone still needs that turned into a real log entry: which piece of equipment, what was done, when, filed against that equipment’s own history rather than as a flat note nobody can search. And every so often, a note buried in otherwise routine language is actually describing something that needs a person’s attention now, not at the next scheduled review.

The pipeline is three steps: extract a fixed record from the note, find the equipment it concerns in a small graph of what the facility has, and let the model decide whether this is routine or needs flagging.

Example run

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Plain-language maintenance log, assembled

Extract a fixed record, walk the graph to the equipment it concerns, and let the model decide between logging it and flagging it.

Level 4 · Tool use
Note arrivesNote arrivesMODELExtract a fixed recordExtract afixed recordWalk the equipment graphWalk theequipment graphMODELLog it, or flag itLog it, or flag itTOOLsave_log_entry(record)save_log_entry(record)TOOLflag_for_review(reason)flag_for_review(reason)Link into equipment historyLink intoequipment historyLog entry savedLog entry saved
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 06Your code chose

A technician's note arrives

"Replaced the belt on pump 3, it was fraying and
starting to smell hot before I shut it down."
0 tokens · 0 ms

Walkthrough

The three steps compose the runnable code already on the structured output, knowledge graphs and function calling pages, unmodified in shape, each written against the site’s own shared examples (structured output reads a warranty record, knowledge graphs walks from a part to its warranty class) rather than a literal maintenance note. Composing them means keeping the same functions and swapping in this job’s own schema and its own two-hop walk (equipment to line, line to last-serviced date), and offering save_log_entry and flag_for_review in place of lookup_part. Nothing here ships that exact pair of tools yet; what’s illustrated above is the shape those three pages’ code already is, run for this job.

The run above shows a note that reads as routine until its last clause. Extraction pulls a clean record (equipment, action, a condition field carrying the technician’s own words) and the graph walk finds pump 3’s line and that it was serviced two months ago, nothing unusual on its own. The model reads the condition field’s “starting to smell hot” alongside a worn belt and calls flag_for_review instead of save_log_entry, the one decided_by: "model" step in the whole run; a note that only said “replaced the belt, routine wear” would have taken the other tool instead, with nothing else in the pipeline changing. Either way your code links the entry into the equipment’s history; what the flag adds is a review task on top of it. Log both tools offered, which one was called, and the graph path behind the record, so a note that should have been flagged and wasn’t can be found afterwards.

What to measure

Knowledge graphs and function calling are each scored on their own by the site’s shared 60-question set (see docs/EVALS.md); structured output is not. None of the three answers this job’s actual task, so build a labeled set of real or synthetic notes instead: the record a person would extract, the equipment it should link to, and whether a facilities lead would flag it. Watch whether the escalation tool gets called on the notes that actually need it, and only those, checked against a sample of routine entries to confirm none should have been flagged instead. Add one measure specific to the graph: how often the named equipment fails to resolve to anything in it, which is a data problem, not a model one. No result file exists for any of the three on this task yet, so this recipe cannot claim a score for any of it.

Variations

  • Add human approval before saving an entry whose equipment id the graph doesn’t recognize: a missing link and a safety concern are different problems and shouldn’t share one gate.
  • Move to a single agent if escalating ever needs more than one lookup: checking recent notes on similar equipment before judging this one.
  • Route by facility area first with routing if the same pipeline serves more than one site, each with its own equipment graph.
  • Rebuild the graph on a schedule, not on every note; knowledge graphs’ own page is direct about the cost: extraction is expensive, and paid once per equipment change, not per entry.

Design choices

Why this level, and when to use another approach

Three techniques compose this recipe. Structured output turns the note into a fixed record (equipment, action, a condition field, a date) validated before anything downstream touches it. Knowledge graphs is what makes “pump 3” mean something: a small graph walk from the named equipment to what line it’s on and when it was last serviced, so the entry lands linked to a real history rather than as an isolated row. Function calling is the one place this recipe climbs past a fixed pipeline: the model is offered two tools (log it routinely, or flag it for review) and decides which one this note actually calls for.

That decision is why level 4, not level 2, is the honest floor here. Structured output and knowledge graphs alone would produce a well-formed, well-linked record every time, but they’d file “smelled hot” with the same routine handling as “replaced the belt”: neither technique reads for risk, only for fields and connections. Function calling’s own page draws the boundary this recipe needs: the model is offered a schema-shaped action and decides on its own whether to reach for it. Structured output’s page draws the same contrast from its own side: at level 4 the model additionally decides whether to use a shape at all, not just what goes in it. A fixed keyword list could catch a few obvious cases, but “starting to smell hot” is exactly the open-ended phrasing such a list can’t be written to catch reliably in advance.

It isn’t worth climbing past that single decision. The model acts once, on one note, with everything it needs already in the extracted record and the graph walk; nothing here asks it to check a result and decide again. A facility that wanted the model to weigh a piece of equipment’s full service history against similar recent notes elsewhere before deciding (several lookups, each depending on the last) is the case a single agent is for, at real added cost for a decision this recipe currently makes in one call.

Composition

Techniques this recipe uses

The highest level it needs is level 4.

Function calling

Sourced

Letting the model call functions that you define.

Structured output

Sourced

Getting answers in a fixed format such as JSON.

Knowledge graphs and GraphRAG

Sourced

Storing facts as entities and relations, for questions that span several documents.

Same shape, other jobs

Pull structured data out of something unstructured

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • Invoices and receipts into an accounting system
  • Key parameters from a datasheet into a parts database
  • An instrument accuracy table into rows per range and per calibration interval
  • A calibration certificate into as-found and as-left readings for a drift record
  • Operator failure notes into cause, location and severity
  • Resumes into a candidate record
  • Lab reports into a results table
  • Log lines into typed events

Last reviewed 09/18/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page