Voice notes into structured entries
Transcribes a voice note, splits it into steps, and turns each step into a structured entry that an eval set checks for accuracy.
SourcedNeeds level 3
A field technician finishes a site visit and dictates a note instead of typing one: “checked the rooftop unit, filter’s fine, belt’s worn and needs replacing before next visit, also the condensate line has a slow drip near the drain pan.” That note needs to become two structured findings a scheduling system can act on, not a paragraph someone reads later and re-types.
Example run
Optional: inspect the implementation trace
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
Voice notes into structured entries
One fixed transcription call, then a fixed split-and-fill sequence. No step chooses what happens next.
The run, step by step
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
The note arrives
audio, 41 seconds. Code calls the transcription model once; this call was never in question.
Walkthrough
The transcript comes back as one block of text. Code splits it on cues a technician’s dictation already tends to carry (“also,” “next,” a pause long enough to mark in the transcript) into a list of candidate findings; this is level 0 work, not a model’s judgment call, the same way document-Q&A’s chunking is just where the source documents already break. Each candidate goes to the model once, with a fixed schema: location, component, condition, an action if one was stated, and a confidence field the model fills honestly rather than guessing when the note didn’t say. A last code step checks every required field came back non-empty, and flags a record that is missing one rather than filing it: the retry-and-validate shape a structured-output step takes anywhere on this site.
Recording is a question to settle before this ships. The technician starts the recording, so nobody is taped unawares, but they should still be told where the audio goes, how long it is kept and who can play it back: a dictated note is a recording of someone’s voice, not only a row in a table. Retention makes the rest easy. If the structured record is what the scheduling system needs, the audio can be deleted once a transcript is accepted, and no archive of voices is left to govern. A note that catches a second voice in the background (a customer talking on site) is a different question, answered differently in different places and one for a lawyer rather than this page; deleting those is the cheap default.
Transcribe locally and the audio never leaves the device; transcribe against a hosted model and every note does. Either way the transcript and the records reach whatever stores site records, which turns a dictated remark into searchable text. Nothing here pauses for approval before a record is filed, so the risk is a misheard word becoming a wrong structured fact nobody checks until the next visit, which is what the eval below is for. Cost per note tracks how long the audio is rather than how many words are in it, since audio is charged by duration on at least one maker’s API (the multimodal page carries the figure); the structuring calls after it are small by comparison.
What to measure
Collect a small set of real-shaped, synthetic dictated notes (a few dozen, covering clean single-finding notes, multi-finding notes, and a few with a misheard-sounding word or a mid-sentence correction) each with a hand-written correct set of records. Measure finding count accuracy (did the split produce the right number of records, not too many or too few), field accuracy per required field against the hand-written answer, and the flag rate: how often a record with a genuinely wrong field also came back with low confidence or a validation failure, versus how often a wrong field slipped through with nothing marking it. That last number is the one worth watching in production: a wrong record nobody is warned about is worse than one the system admits it isn’t sure of. No run of this has been scored yet.
Variations
- Split transcription and structuring across two models (a small, fast speech-to-text one for the transcript, a larger one only for the structuring pass) if a local model can transcribe well enough to skip sending audio anywhere.
- Move to a single agent once technicians start dictating messier notes than a fixed splitter can reliably segment.
- Add human approval for any record whose confidence field or validation came back low, instead of filing it straight through.
- Delete the audio as soon as a transcript is accepted, keeping only the text and the records, so the system stops holding recordings of people’s voices at all.
Design choices
Why this level, and when to use another approach
Multimodal covers turning the voice note into text: one request that carries the audio itself and comes back with a transcript: level 1, the same one-call shape as any other request, with audio where a paragraph would be. Prompt chaining is the technique doing the real work after that: split the transcript into individual findings, in the order the technician said them, and fill a fixed record (location, component, condition, action needed) for each one. Evals is how “did it get this right” is measured for a job with no citation to check against, the way document-Q&A has.
Both composed techniques put every decision in code. Transcribing is one fixed call: the code
always makes it, once, and never asks the model whether to. Splitting the transcript into findings
and filling each one’s record is a fixed sequence too: every step runs in the same order
regardless of what came back from the last one. Set that against the site’s own thesis: this
recipe reaches level 3 on the ladder, and yet not one step in its run is decided_by: "model".
The model only fills in what a fixed step already decided to ask it for.
Climbing to a single agent would trade that predictability for judgment the fixed order can’t offer: a note where findings interrupt each other (“actually, before the belt, check the filter too”), get corrected mid-sentence, or describe two sites in one recording. A technician who dictates cleanly, one finding after another, the way this recipe’s illustrated run does, doesn’t need a model deciding how to segment the note; a fixed splitter does the same job for less and never disagrees with itself about where one finding ends and the next begins.
Techniques this recipe uses
The highest level it needs is level 3.
Images, audio and video
SourcedGiving the model images, audio, video and documents, and getting them back.
Pull structured data out of something unstructured
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- Invoices and receipts into an accounting system
- Key parameters from a datasheet into a parts database
- An instrument accuracy table into rows per range and per calibration interval
- A calibration certificate into as-found and as-left readings for a drift record
- Operator failure notes into cause, location and severity
- Resumes into a candidate record
- Lab reports into a results table
- Log lines into typed events
Last reviewed 09/18/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page