# Voice agents

_Level 05 · Agent loops · sourced_

Agents you talk to in real time.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a spoken interaction through interpretation, clarification, and a response or action. Focus on names, quantities, interruptions, and corrections that can change the user's intent.

**Assumptions:** This walkthrough represents speech as text. A real voice system must handle audio uncertainty and turn-taking as well as the task itself.

**Design choices:** Confirm consequential or easily misheard details, while letting ordinary conversation remain natural. Choose whether a correction replaces a draft or arrives after a commitment.

**Request:** Book a workshop place and confirm details first.

**Starting evidence:** Transcript fixture: Saturday at ten. Slots: 10 am or 10 pm. Name heard as Lee or Leigh.

**Action and control:** Clarify time and spelling before confirmation. This transcript fixture does not process audio or make a booking.

**Stage records (authored, not executed):**

### Input record

Transcript fixture: Saturday at ten. Slots: 10 am or 10 pm. Name heard as Lee or Leigh.

What changed: Establish the facts supplied for this version of the task.

### Design note

Confirm consequential or easily misheard details, while letting ordinary conversation remain natural. Choose whether a correction replaces a draft or arrives after a commitment.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Clarify time and spelling before confirmation. This transcript fixture does not process audio or make a booking.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Proposal: Saturday 10 am, Leigh. Wait for confirmation before any booking action.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Transcript/audio controls, turn state, correction, scoped confirmation, and a clearly simulated booking result.

If the result falls short:
If speech is unclear or the user interrupts, preserve the last confirmed intent and clarify the disputed part. Check action state before attempting a second submission.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for scheduling, assistance, or hands-free workflows. Adapt confirmation to the action's consequence; do not turn every spoken sentence into an approval ceremony.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Proposal: Saturday 10 am, Leigh. Wait for confirmation before any booking action.

**Change something — User interrupts with Sunday instead:** Stop the outdated confirmation, check the new date, and reconfirm; do not commit the old slot.

**Decision:** Should an interrupted confirmation trigger booking?

**Answer:** No; resolve the changed request first.

**Why:** Handle interruptions, an ambiguous date, and a misunderstood name; confirm before committing the booking.

**Review criteria:** Transcript/audio controls, turn state, correction, scoped confirmation, and a clearly simulated booking result.

**Recovery:** If speech is unclear or the user interrupts, preserve the last confirmed intent and clarify the disputed part. Check action state before attempting a second submission.

**Adapt it:** Use this for scheduling, assistance, or hands-free workflows. Adapt confirmation to the action's consequence; do not turn every spoken sentence into an approval ceremony.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a spoken interaction through interpretation, clarification, and a response or action. Focus on names, quantities, interruptions, and corrections that can change the user's intent.

**Assumptions:** This walkthrough represents speech as text. A real voice system must handle audio uncertainty and turn-taking as well as the task itself.

**Design choices:** Confirm consequential or easily misheard details, while letting ordinary conversation remain natural. Choose whether a correction replaces a draft or arrives after a commitment.

**Request:** Read back a proposed configuration while the engineer works hands-free.

**Starting evidence:** Transcript fixture: set channel A to two point zero volts. Recognition could confuse two with twenty. No hardware tool is enabled.

**Action and control:** Read back value, unit, and channel and request explicit confirmation; spoken understanding is not execution authority.

**Stage records (authored, not executed):**

### Input record

Transcript fixture: set channel A to two point zero volts. Recognition could confuse two with twenty. No hardware tool is enabled.

What changed: Establish the facts supplied for this version of the task.

### Design note

Confirm consequential or easily misheard details, while letting ordinary conversation remain natural. Choose whether a correction replaces a draft or arrives after a commitment.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Read back value, unit, and channel and request explicit confirmation; spoken understanding is not execution authority.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Proposal: channel A, 2.0 V, not applied. Engineer confirms or corrects the transcript before any separately authorized action.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Inspect the transcript, interruption, corrected readback, and absence of instrument execution.

If the result falls short:
If speech is unclear or the user interrupts, preserve the last confirmed intent and clarify the disputed part. Check action state before attempting a second submission.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for scheduling, assistance, or hands-free workflows. Adapt confirmation to the action's consequence; do not turn every spoken sentence into an approval ceremony.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Proposal: channel A, 2.0 V, not applied. Engineer confirms or corrects the transcript before any separately authorized action.

**Change something — Engineer interrupts with no, channel C:** Cancel the old proposal and restate the complete corrected configuration. Do not execute A from a partial confirmation.

**Decision:** Should an interrupted set-point proposal remain actionable?

**Answer:** No; invalidate and reconfirm the corrected proposal.

**Why:** Voice timing and recognition errors require explicit state handling, particularly around consequential actions.

**Review criteria:** Inspect the transcript, interruption, corrected readback, and absence of instrument execution.

**Recovery:** If speech is unclear or the user interrupts, preserve the last confirmed intent and clarify the disputed part. Check action state before attempting a second submission.

**Adapt it:** Use this for scheduling, assistance, or hands-free workflows. Adapt confirmation to the action's consequence; do not turn every spoken sentence into an approval ceremony.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a spoken interaction through interpretation, clarification, and a response or action. Focus on names, quantities, interruptions, and corrections that can change the user's intent.

**Assumptions:** This walkthrough represents speech as text. A real voice system must handle audio uncertainty and turn-taking as well as the task itself.

**Design choices:** Confirm consequential or easily misheard details, while letting ordinary conversation remain natural. Choose whether a correction replaces a draft or arrives after a commitment.

**Request:** Collect a spoken project update and draft it for the weekly report.

**Starting evidence:** Transcript: we expect completion Friday, unless the supplier slips. The statement is a forecast.

**Action and control:** Preserve uncertainty and read back the update before incorporating it into the report draft.

**Stage records (authored, not executed):**

### Input record

Transcript: we expect completion Friday, unless the supplier slips. The statement is a forecast.

What changed: Establish the facts supplied for this version of the task.

### Design note

Confirm consequential or easily misheard details, while letting ordinary conversation remain natural. Choose whether a correction replaces a draft or arrives after a commitment.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Preserve uncertainty and read back the update before incorporating it into the report draft.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Draft: completion forecast Friday, contingent on supplier delivery. Await owner confirmation; do not send.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Verify date, forecast versus commitment, owner confirmation, and report version.

If the result falls short:
If speech is unclear or the user interrupts, preserve the last confirmed intent and clarify the disputed part. Check action state before attempting a second submission.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for scheduling, assistance, or hands-free workflows. Adapt confirmation to the action's consequence; do not turn every spoken sentence into an approval ceremony.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Draft: completion forecast Friday, contingent on supplier delivery. Await owner confirmation; do not send.

**Change something — User interrupts with next Friday, not this Friday:** Invalidate the earlier date and clarify the calendar date within the reporting period.

**Decision:** Should the first recognized date survive a spoken correction?

**Answer:** No; resolve and confirm the corrected date.

**Why:** Voice interaction must track corrections and qualifications, not just transcribe fluent sentences.

**Review criteria:** Verify date, forecast versus commitment, owner confirmation, and report version.

**Recovery:** If speech is unclear or the user interrupts, preserve the last confirmed intent and clarify the disputed part. Check action state before attempting a second submission.

**Adapt it:** Use this for scheduling, assistance, or hands-free workflows. Adapt confirmation to the action's consequence; do not turn every spoken sentence into an approval ceremony.

A voice agent is a [single agent](/gradient_ascent/techniques/single-agent/) you talk to instead
of type to: speech in, speech out, in real time. Real time adds a decision chat does not need:
something has to decide when the caller has stopped talking, and what happens if they start again
while the agent is still going. OpenAI's Realtime API documents voice activity detection that
"will determine when the user has started or stopped speaking and respond automatically," and for
interruption: "The server will automatically truncate unplayed audio when there's a user
interruption"[1]. Google calls the same behavior "Barge-in": "Users can interrupt the
model at any time for responsive interactions"[2].

The example below is a text simulation, not a voice pipeline: nothing here records, streams,
transcribes or synthesizes sound. What it can honestly show is the control flow (which choices
are the model's and which belong to the system around it) using `Message` content parts to
reference audio without pretending to process it.

Level 5 still means the model decides the action and the stop: here, whether to keep talking or
yield the floor. Your code enforces the caps: a latency budget on a turn, and an interruption
that is never the model's own choice to make.

This page is sourced, not measured: every latency and turn-taking claim below comes from a
maker's own documentation, and no call has been recorded and scored here. It is illustrated.

_The web page for this technique includes an interactive step-through of Level 5 · Voice agents. The same steps are described in the sections below._

## Practical guidance

If you are setting one up for your business rather than just talking to one, two things have to be
in place before it takes a single real call: the disclosure, and the way out to a person.

Say plainly, before the caller can say anything else, that they are talking to an AI. ElevenLabs
requires exactly this of businesses using its agents: telling callers "They are interacting with AI
rather than a human" and that "Their conversations are being recorded and may be shared with
ElevenLabs and its third-party large language model providers,"[4] with that notice
"presented immediately prior to any interaction"[4], not after the first exchange.
ElevenLabs is explicit that meeting this is the organization's own responsibility, not legal
advice[4], and what the law actually requires varies by place; check it for where you
operate rather than assuming a platform's minimum covers you.

Build the way out next: a phrase that reliably hands the call to a person ("talk to someone," "I
need a human") and a real person or queue on the other end of it, not a dead end. Write the script
for that handoff the same way you would write the disclosure: plainly, and tested before it goes
live.

Then call it yourself, on your own phone, before a customer does. Interrupt it mid-sentence on
purpose and check it actually stops: LiveKit's own documentation describes a framework that "pauses
the agent's speech whenever it detects user speech in the input audio"[3], and a real
deployment should behave the same way, not talk over you. Say a stray "mm-hmm" while it is
mid-answer and check it keeps going rather than restarting, since a good one tells "true
interruptions from conversational backchanneling" apart[3]. Ask for the handoff phrase
and confirm it actually reaches your fallback, not a dropped call. Time the pause between finishing
a sentence and its reply: a long silence with nothing said to fill it reads as broken, not
thoughtful.

If what callers ask is small and known in advance (hours, an address, a balance) a script or a menu
answers it without needing a conversation to manage at all.

## Implementation details

The example simulates one turn's control flow with a text model and no audio anywhere. The caller's
utterance is carried as an `AudioPart` label next to a `TextPart` transcript: the transcript is
what the model actually reads, and the label is what a trace shows in place of sound it never
processed, exactly the pattern `examples/common/model.py` documents for a run that never reaches a
real speech model.

The model generates its answer in chunks. After each one it either calls `continue_speaking`,
meaning it has more to say, or calls nothing, meaning it is done and the floor returns to the
caller: `decided_by: "model"` either way, the same shape every level-5 example on this site uses
for its stop. `MAX_CHUNKS_PER_TURN` (3) stands in for a latency budget: past a certain number of
chunks a real system has to cut the agent off to stay responsive, whatever the model would have
said next. `interrupt_after_chunk` simulates a caller starting to talk mid-turn; when it fires, the
code cuts the agent off immediately, without asking the model anything: `decided_by: "code"`,
because a real interruption is a signal the system acts on the instant it arrives, not a choice the
model gets to weigh in on. OpenAI's own guidance for evaluating a voice agent puts this on the
list of things to measure: "audible response timing, unwanted silence, overlap, and yielding to
interruptions"[5].

`examples/voice_agents/run.py` (lines 38-86)

```python
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    interrupt_after_chunk: int | None = None,
    max_chunks: int = MAX_CHUNKS_PER_TURN,
    max_tokens: int = MAX_TOKENS,
) -> Answer:
    del embedder  # no retrieval here; this page is about turn control, not what gets said
    content = [AudioPart(media_type="audio/wav", label=question), TextPart(text=question)]
    tracer.record(kind="code", decided_by="code", title="Caller speaks", detail=f"[audio] {question[:150]}")
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=content)]

    chunks: list[str] = []
    tokens_used = 0
    for chunk_index in range(max_chunks):
        if interrupt_after_chunk is not None and chunk_index == interrupt_after_chunk:
            tracer.record(
                kind="code", decided_by="code", title="Caller interrupts; code cuts the agent off",
                detail=f"stopped after {len(chunks)} chunk(s)",
            )
            break

        completion = model.complete(messages, tools=[CONTINUE_TOOL], max_tokens=100)
        tokens_used += completion.tokens_in + completion.tokens_out
        chunks.append(completion.text)

        if not completion.tool_calls:
            record_completion(tracer, decided_by="model", title="Model finishes and yields the floor", completion=completion)
            break

        record_completion(tracer, decided_by="model", title="Model chooses to keep talking", completion=completion)
        messages.append(Message(role="assistant", content=completion.text))

        if tokens_used >= max_tokens:
            tracer.record(
                kind="code", decided_by="code", title="Latency budget reached; code cuts the agent off",
                detail=f"{tokens_used} >= {max_tokens} tokens",
            )
            break
    else:
        tracer.record(
            kind="code", decided_by="code", title="Chunk cap reached; code cuts the agent off",
            detail=f"{max_chunks} chunks",
        )

    return Answer(text=" ".join(c for c in chunks if c), citations=[])
```

The same control flow answers a hands-busy question at a bench: an engineer with both hands full
asks for the last reading and expects to hear it back, not a new one. That is honest only if the
voice layer never produces the number itself. `SYSTEM_PROMPT` here is about turn control, not
what gets said, so a bench version would have to add one more rule: read back a value the caller
is given, never estimate, recall or restate one from anywhere else.

Run it yourself:

`examples/voice_agents/README.md` (lines 18-18)

```text
python -m examples.voice_agents --model stub:scripted
```

## When you do not need this

Try plain [chat](/gradient_ascent/techniques/chat/) first if the interaction does not actually
need to happen in real time: turn-taking, interruption and latency budgets are all cost you pay
for synchrony you may not need.

Try a fixed workflow instead (transcribe the utterance, then
[route](/gradient_ascent/techniques/routing/) it to one of a few known intents, then answer with a
scripted response) if the things a caller might ask are small and known in advance; that is
cheaper and does not need the model to manage the turn itself.

Move up to a voice agent once the range of things a caller might say is too open for a fixed set of
intents, and the conversation genuinely needs to flow rather than follow a menu.

## Failure modes

### The agent talks over the caller

- **How to notice it:** The caller starts speaking and the agent keeps going instead of yielding immediately, breaking the sense that anyone is actually listening.
- **How to test for it:** Interrupt mid-sentence and time how long the agent keeps talking before it stops. LiveKit documents its framework pausing agent speech the instant it detects caller speech; a noticeable delay past that is this failure.

### A false interruption derails the agent

- **How to notice it:** A stray "mm-hmm" or a cough gets read as a real interruption, and the agent restarts or drops what it was saying instead of continuing.
- **How to test for it:** Say a short backchannel sound while the agent is mid-answer and check whether it resumes from where it left off, which is the recovery LiveKit's documentation describes, or restarts from scratch.

### No disclosure, or disclosure too late

- **How to notice it:** The caller is well into the conversation before anything tells them they are talking to an AI, if anything ever does.
- **How to test for it:** Start a session and check whether a disclosure plays before you can say anything at all. ElevenLabs requires this to be presented immediately prior to any interaction, not after the first exchange.

### The latency budget is blown silently

- **How to notice it:** A turn takes long enough that the pause reads as dead air, with nothing telling the caller the agent is still working.
- **How to test for it:** Measure the gap between the caller finishing and the agent's first audible response across many turns, not just once; an occasional slow turn with no filler or acknowledgment is this failure even if the average looks fine.

### The cap cuts off mid-sentence with no recovery

- **How to notice it:** A chunk or token cap ends the turn partway through a sentence, and the agent neither finishes the thought nor says anything to cover the abrupt stop.
- **How to test for it:** Force a low chunk cap on an answer that needs more than one chunk (this page's own test suite does exactly this) and check whether what comes back reads as a real, if short, answer or as speech cut off mid-word.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one-chunk turn:** 1
- **Model calls, full chunk cap reached:** 3
- **Tokens in, one chunk:** ~180–260
- **Wall time, one chunk (illustrated):** ~0.4s

**Compared with a text chat turn (level 1).** A real voice turn adds speech-to-text and text-to-speech latency on top of whatever this diagram shows, neither of which this text-only example measures; the illustrated numbers cover only the control-flow decisions themselves.

## How to Evaluate It

`voice_agents` does not do the site's own question-answering task and is not part of the shared
60-question set: a text-simulated turn loop has no document to cite. What changes for evaluation
here is the task itself: voice needs its own small eval, on its own page, the way this site treats
every technique whose running task is not a fair test of it. What that eval would measure is
specific to a spoken conversation: yielding to interruptions (does the agent actually stop when
talked over), time from the caller finishing to the agent's first audible response, and
false-interrupt recovery (does a backchannel derail it). `scripts/eval_run.py` knows
`voice_agents` and refuses to score it, printing that reason; no runner for the measures above
exists yet.

## Run it

**What to monitor.** How quickly the agent actually stops when talked over, the false-interrupt rate, and response latency per turn, tracked separately from whatever the underlying model's own latency is.

**Cost at volume.** A real deployment pays for speech-to-text and text-to-speech on top of every model call this example's control flow makes, often the larger share of per-minute cost at volume, not the model itself.

**How it fails in production.** A change to the underlying model shifts its per-chunk timing enough that the fixed latency budget starts cutting off answers that used to finish in time, with nothing in a dashboard built around token counts likely to show it.

**What to log.** Every chunk generated, the time between them, whether a turn ended by the model yielding, an interruption, or a cap, and whether whatever disclosure your platform or your lawyers require actually played before the conversation started.

## Try it

1. **Use it.** Start a conversation with a voice assistant and interrupt it mid-sentence on purpose. Does it stop immediately? Does anything at the start of the call tell you it is an AI?
2. **Build it.** Run python -m examples.voice_agents --model stub:scripted from the repo root. The model speaks one chunk, keeps the floor, speaks a second, and yields on its own: two decisions, not one reply. Now give every entry in SCRIPTED (examples/voice_agents/__main__.py) a continue_speaking call and add a third. Nothing yields, so the code cuts the agent off at the cap.
3. **Either lane.** Write the disclosure sentence you would want at the start of a call with an AI. Then check the SYSTEM_PROMPT in examples/voice_agents/run.py: it does nothing like it, on purpose, since this example is about turn control. Where would you add it?


## Sources

1. [Realtime conversations](https://developers.openai.com/api/docs/guides/realtime-conversations) — OpenAI (API documentation) (accessed 2026-09-19)
2. [Gemini Live API overview (archived copy)](https://web.archive.org/web/20260915171334id_/https://ai.google.dev/gemini-api/docs/live-api) — Google (Gemini API documentation, via the Internet Archive) (accessed 2026-09-19)
3. [Turns overview](https://docs.livekit.io/agents/logic/turns/) — LiveKit (accessed 2026-09-19)
4. [Disclosure requirements](https://elevenlabs.io/docs/eleven-agents/legal/disclosure-requirement) — ElevenLabs (accessed 2026-09-19)
5. [Voice agents](https://developers.openai.com/api/docs/guides/voice-agents) — OpenAI (API documentation) (accessed 2026-09-19)


Last reviewed 2026-09-19.
