Recipe

Turn a meeting transcript into decisions and owners

Turn a transcript into decisions, owners, and open questions in one model call. Someone who attended reviews the draft before it is shared.

SourcedNeeds level 1

Someone runs a weekly team meeting and has a transcript afterward: a call recording turned to text, or the notes someone typed live while people talked. What they want back is short. Which decisions actually got made. Who owns each one. When it is due, if anyone said. What is still open. Not a summary of everything that was discussed, a log somebody can act on without rereading the whole meeting to find the four sentences that mattered.

The transcript already has everything this job needs. Nobody has to look anything up, check a policy, or compare against last week’s minutes; the words on the page are the whole source. That is what keeps this at level 1: one call reads the transcript once and returns a fixed shape, and the shape is the entire value this recipe adds over reading the transcript and typing the same four things by hand.

Not in scope: filing a decision into a task tracker, checking it against what a past meeting decided, or judging whether the team decided the right thing. And this recipe cannot tell a reader whether something the model calls a decision was actually agreed to in the room. Only a person who was there can say that, which is why one reads the notes before they go anywhere.

Example run

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Meeting notes, assembled

Ask for decisions in a fixed shape, then check every quote against the transcript before a person reads it.

Level 1 · Direct prompting
Transcript arrivesTranscript arrivesMODELAsk the model for JSONAsk the model for JSONValidate against the schemaValidate againstthe schemaCheck each quote against the transcriptCheck each quoteagainst the transcriptPERSONPerson who was in the room reads itPerson who was inthe room reads itNotes acted onNotes acted onTranscript arrivesTranscript arrivesMODELAsk the model for JSONAsk the model for JSONValidate against the schemaValidate againstthe schemaCheck each quote against the transcriptCheck each quoteagainst the transcriptPERSONPerson who was in the room reads itPerson who was inthe room reads itNotes acted onNotes acted on
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 05Your code chose

The transcript arrives

Weekly sync, Thornwood product team: 25 lines, 4 attendees,
a shipped decision, a deferred discount, a role-owned action
item, and one open question about a vendor's pricing tier
0 tokens · 0 ms

Walkthrough

Run against the sample transcript in examples/meeting_notes/run.py, a 25-line invented weekly sync with four attendees, the model reads the whole thing once and replies with attendees, decisions and open questions in the fixed shape. Two of the three decisions it proposes survive the code that follows: shipping the new signup flow on Friday, owned by the person who said they would cut the release, and the FAQ for that flow going to whoever is on support rotation next week, an owner named by role rather than by name because that is genuinely how the room left it. The one open question, whether a vendor’s new pricing tier applies to this team, comes back alongside them.

View code: run
examples/meeting_notes/run.py · lines 196–238
def run(transcript: str, model: Model, tracer: Tracer, *, max_tokens: int = 700) -> MeetingNotes:
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=transcript)]
    record: dict = {}
    problems: list[str] = []
    for attempt in range(MAX_RETRIES + 1):
        completion = model.complete(messages, schema=SCHEMA, max_tokens=max_tokens)
        tracer.record(
            kind="model",
            decided_by="code",
            title="Ask the model for JSON" if attempt == 0 else "Ask again with the validation error",
            detail=completion.text[:200],
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
        try:
            record = json.loads(completion.text)
            problems = _validate(record)
        except json.JSONDecodeError as exc:
            record, problems = {}, [f"invalid JSON: {exc}"]
        tracer.record(kind="code", decided_by="code", title="Validate against the schema", detail="; ".join(problems) or "valid")
        if not problems:
            break
        if attempt < MAX_RETRIES:
            messages.append(Message(role="user", content=f"That did not validate: {'; '.join(problems)}. Reply again with corrected JSON only."))

    if problems:
        tracer.record(kind="code", decided_by="code", title="Give up after the retry", detail="; ".join(problems))
        return MeetingNotes(attendees=(), decisions=(), dropped=(), open_questions=(), error="; ".join(problems))

    kept, dropped = _check_quotes(record["decisions"], transcript)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check each decision's quote against the transcript",
        detail=f"{len(kept)} kept, {len(dropped)} dropped" + (f": {'; '.join(d.decision for d in dropped)}" if dropped else ""),
    )
    return MeetingNotes(
        attendees=tuple(record["attendees"]),
        decisions=tuple(kept),
        dropped=tuple(dropped),
        open_questions=tuple(record["open_questions"]),
    )

The fourth decision in that run, a marketing budget cut nobody in this transcript raised, is there on purpose: its quote does not appear anywhere in the source text. The quote check is a plain substring comparison after both strings are stripped down to single spaces, since a model’s reply often re-wraps a sentence’s line breaks even when it copies the words correctly.

View code: check quotes
examples/meeting_notes/run.py · lines 183–193
def _check_quotes(decisions: list[dict], transcript: str) -> tuple[list[Decision], list[DroppedDecision]]:
    haystack = _normalize_ws(transcript)
    kept: list[Decision] = []
    dropped: list[DroppedDecision] = []
    for d in decisions:
        quote = _normalize_ws(d["quote"])
        if quote and quote in haystack:
            kept.append(Decision(decision=d["decision"], owner=d["owner"], due_date=d["due_date"], quote=d["quote"]))
        else:
            dropped.append(DroppedDecision(decision=d["decision"], quote=d["quote"], reason="quote not found in transcript"))
    return kept, dropped

Read against the real transcript, that run keeps two decisions and drops one, reported by name rather than removed, at a cost of 712 tokens in and 194 tokens out for the single call. A reply that fails validation the first time, tested separately, costs a second call and doubles the token count for that meeting; the retry exists so a stray formatting mistake does not sink the whole run, not to try harder at reading the meeting.

What it costs

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1, or 2 if the first reply fails validationModel calls, one meeting
712Tokens in, this run
194Tokens out, this run
2 kept, 1 droppedDecisions kept and dropped, this run
Compared with a second pass that checks every kept decision against the transcript (level 3, write and check)Adding a check-and-retry pass over the decisions this run kept would roughly double the call count for a meeting this size, the same shape climbing to a write-and-check workflow costs everywhere else on this site. It earns that cost once real runs show the single call keeping something it should not, not before: the quote check above already catches the failure a substring comparison can catch.

The unit here is per meeting, not per month. A team running this once after every weekly sync spends under a thousand tokens to get four names and two dates typed out in a shape somebody can act on, instead of a person rereading a transcript they would otherwise have to open anyway. Volume does not change this recipe’s economics the way it changes a production line’s: there is one transcript, one call, and one person reading the result before anything happens.

How it fails

A model that summarizes instead of extracting

How to notice it
The reply is valid JSON and every decision has a real quote behind it, but reading the list feels like reading a recap of the meeting rather than a list of things that got settled: entries with no clear owner, or a due date left blank on something that plainly needed one.
How to test for it
Read the kept decisions against the transcript by hand on a sample of real runs. Nothing in the schema or the quote check can tell a summarized talking point from an actual decision; both can carry an accurate quote.

An owner invented from context

How to notice it
A decision names a person as its owner who was never actually said to own it, guessed from who talked about the topic most rather than read off what the room actually agreed.
How to test for it
tests/test_example_meeting_notes.py checks the other half of this directly: when the transcript never names an owner, the code accepts the literal word unassigned unchanged rather than filling in a name of its own. Attack it by scripting a reply where the model invents a plausible-looking name for an unnamed action item and confirm nothing downstream of run catches that on its own; only a person comparing the owner to the transcript can.

A decision that was discussed but never made

How to notice it
Something the room explicitly put off, or debated without agreeing on, comes back in the decisions list anyway, backed by a real sentence from the part of the meeting where people were still arguing about it.
How to test for it
tests/test_example_meeting_notes.py scripts exactly this: a reply that claims the team decided to raise a discount, quoting, word for word, the sentence where one attendee asked for another week before anyone committed to a number. The quote check keeps the decision, since the words really are in the transcript, which is the limit of what a substring comparison can catch.

What to measure

A right answer is a kept-decisions list that matches what a person who was in the meeting would also call a decision, each with the owner and due date that person would write down themselves and a quote that actually supports it, with nothing real left off the list. Collect ten to fifteen of the reader’s own transcripts with a person’s own notes written alongside them before tuning anything; one team’s meetings vary enough in how people phrase agreement that five is too few to trust a rate from.

The confusion that matters is not spread evenly across the four fields. It is calling something a decision that was not one: the deferred-item test above shows a wrong decision can carry a real, accurate quote, so nothing in the pipeline catches it before a person does. A real decision the recipe misses costs a rereading of the transcript to find it. A decision it invents can cost the thing itself, once someone reads the notes instead of the meeting and acts on what they say. Score false decisions, not missed ones, as the number to drive toward zero.

No result file exists for this recipe yet, so it claims no score. What exists is the shape of the check a person should run by hand, which is the same one the failure modes above describe.

Variations

  • Swap the schema for a different meeting shape, a one-on-one or a board vote, and keep the ask-validate-retry contract and the quote check exactly as written; only the fields change.
  • Move to write and check at level 3 once real runs show the single call keeping a decision the quote check cannot catch, or missing one a person would have caught. Add a second pass that checks each kept decision against the rule the failure modes above describe, rather than a longer prompt asking the first call to be more careful.
  • Move to retrieval at level 2 only once these notes have to be checked against, or linked to, what past meetings decided. One transcript needs none of that.
  • Start from a recording instead of a transcript. Feeding audio in directly replaces the first step; the schema, the validation and the quote check downstream of it do not change.

Design choices

Why this level, and when to use another approach

Two techniques compose this recipe. Prompt engineering is the system prompt itself: a required shape instead of prose, an explicit line about what counts as a decision (something the room agreed to, not something raised or put off), and instructions for the two fields a model would otherwise be tempted to guess, the owner and the due date. Structured output is the schema, the validation and the one retry: every decision comes back as an object with the same four fields, code checks it against the schema before accepting anything, and a reply still invalid after one retry is reported as failed rather than returned as though it worked.

One more check runs entirely in code, level 0, and it is the one this page is really about: whether a decision’s quote is actually a sentence from the transcript. That needs no model at all, only a whitespace-normalized substring check, and it is the difference between a decision a person can verify at a glance and one they have to take on faith.

Level 2 would add retrieval: a search over other meetings’ notes before answering, so a decision could be checked against what this team agreed to last time, or linked to the one it changes. Nothing in a single transcript needs that. It would cost a second call, an index of every past meeting’s notes to search, and a new way to be wrong, the wrong past meeting retrieved. It is worth adding once these notes are read against a history, not when they are written from one transcript.

Level 0 on its own is not enough: which sentences in a transcript describe something the room actually settled, as opposed to a proposal, a question, or an idea two people talked themselves out of, is not a rule a regular expression can write. Speakers phrase agreement a dozen different ways in one meeting, and reading which one it was this time is the part of the job that needs a model.

Composition

Techniques this recipe uses

The highest level it needs is level 1.

Prompt engineering

Sourced

Writing instructions that get consistent results.

Structured output

Sourced

Getting answers in a fixed format such as JSON.

Same shape, other jobs

Turn one piece of text into another

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • Summarize a meeting transcript
  • Explain a compiler error or a stack trace
  • Write release notes from a list of commits
  • Rewrite a test procedure for a less experienced operator
  • Write a characterization report around numbers that are already computed
  • Translate a supplier's datasheet excerpt
  • Turn bullet points into a status report

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page