# Voice notes into structured entries

_Recipe · needs level 3_

Transcribes a voice note, splits it into steps, and turns each step into a structured entry that an eval set checks for accuracy.


A field technician finishes a site visit and dictates a note instead of typing one: "checked the
rooftop unit, filter's fine, belt's worn and needs replacing before next visit, also the
condensate line has a slow drip near the drain pan." That note needs to become two structured
findings a scheduling system can act on, not a paragraph someone reads later and re-types.

## Example run

_The web page for this technique includes an interactive step-through of Level 3 · assembled for this recipe. The same steps are described in the sections below._

## Walkthrough

The transcript comes back as one block of text. Code splits it on cues a technician's dictation
already tends to carry ("also," "next," a pause long enough to mark in the transcript) into a
list of candidate findings; this is level 0 work, not a model's judgment call, the same way
document-Q&A's chunking is just where the source documents already break. Each candidate goes to
the model once, with a fixed schema: location, component, condition, an action if one was stated,
and a confidence field the model fills honestly rather than guessing when the note didn't say. A
last code step checks every required field came back non-empty, and flags a record that is
missing one rather than filing it: the retry-and-validate shape a structured-output step takes
anywhere on this site.

Recording is a question to settle before this ships. The technician starts the recording, so
nobody is taped unawares, but they should still be told where the audio goes, how long it is kept
and who can play it back: a dictated note is a recording of someone's voice, not only a row in a
table. Retention makes the rest easy. If the structured record is what the scheduling system
needs, the audio can be deleted once a transcript is accepted, and no archive of voices is left to
govern. A note that catches a second voice in the background (a customer talking on site) is a
different question, answered differently in different places and one for a lawyer rather than this
page; deleting those is the cheap default.

Transcribe locally and the audio never leaves the device; transcribe against a hosted model and
every note does. Either way the transcript and the records reach whatever stores site records,
which turns a dictated remark into searchable text. Nothing here pauses for approval before a
record is filed, so the risk is a misheard word becoming a wrong structured fact nobody checks
until the next visit, which is what the eval below is for. Cost per note tracks how long the
audio is rather than how many words are in it, since audio is charged by duration on at least one
maker's API (the multimodal page carries the figure); the structuring calls after it are small by
comparison.

## What to measure

Collect a small set of real-shaped, synthetic dictated notes (a few dozen, covering clean
single-finding notes, multi-finding notes, and a few with a misheard-sounding word or a
mid-sentence correction) each with a hand-written correct set of records. Measure **finding
count accuracy** (did the split produce the right number of records, not too many or too few),
**field accuracy** per required field against the hand-written answer, and the **flag rate**: how
often a record with a genuinely wrong field also came back with low confidence or a validation
failure, versus how often a wrong field slipped through with nothing marking it. That last number
is the one worth watching in production: a wrong record nobody is warned about is worse than one
the system admits it isn't sure of. No run of this has been scored yet.

## Variations

- Split transcription and structuring across two models (a small, fast speech-to-text one for
  the transcript, a larger one only for the structuring pass) if a local model can transcribe
  well enough to skip sending audio anywhere.
- Move to [a single agent](/gradient_ascent/techniques/single-agent/) once technicians start
  dictating messier notes than a fixed splitter can reliably segment.
- Add [human approval](/gradient_ascent/techniques/human-in-the-loop/) for any record whose
  confidence field or validation came back low, instead of filing it straight through.
- Delete the audio as soon as a transcript is accepted, keeping only the text and the records, so
  the system stops holding recordings of people's voices at all.
## Design choices

### Why this level, and when to use another approach

[Multimodal](/gradient_ascent/techniques/multimodal/) covers turning the voice note into text:
one request that carries the audio itself and comes back with a transcript: level 1, the same
one-call shape as any other request, with audio where a paragraph would be.
[Prompt
chaining](/gradient_ascent/techniques/prompt-chaining/) is the technique doing the real work after that: split the transcript into
individual findings, in the order the technician said them, and fill a fixed record (location,
component, condition, action needed) for each one. [Evals](/gradient_ascent/techniques/evals/)
is how "did it get this right" is measured for a job with no citation to check against, the way
document-Q&A has.

Both composed techniques put every decision in code. Transcribing is one fixed call: the code
always makes it, once, and never asks the model whether to. Splitting the transcript into findings
and filling each one's record is a fixed sequence too: every step runs in the same order
regardless of what came back from the last one. Set that against the site's own thesis: this
recipe reaches level 3 on the ladder, and yet not one step in its run is `decided_by: "model"`.
The model only fills in what a fixed step already decided to ask it for.

Climbing to [a single agent](/gradient_ascent/techniques/single-agent/) would trade that
predictability for judgment the fixed order can't offer: a note where findings interrupt each
other ("actually, before the belt, check the filter too"), get corrected mid-sentence, or
describe two sites in one recording. A technician who dictates cleanly, one finding after
another, the way this recipe's illustrated run does, doesn't need a model deciding how to segment
the note; a fixed splitter does the same job for less and never disagrees with itself about where
one finding ends and the next begins.



Last reviewed 2026-09-18.
