Level 05 · Agent loops

Voice agents

Agents you talk to in real time.

Sourced

Concept at a glance

Listen and respond without losing the turn.

Feedback loopConceptual illustration
Listen and respond without losing the turn.Spoken input leads to Agent. Agent leads to Speak. Speak leads to Listen again. Listen again leads to Agent as feedback. Conversation timing and interruption handling surround the model’s response.Spoken inputA person starts a turnAgentUnderstand and respondSpeakDeliver the responseListen againNew turn or interruptionListen and respond without losing the turn.Spoken input leads to Agent. Agent leads to Speak. Speak leads to Listen again. Listen again leads to Agent as feedback. Conversation timing and interruption handling surround the model’s response.Spoken inputA person starts a turnAgentUnderstand and respondSpeakDeliver the responseListen againNew turn or interruption

Ending or continuingListen for a new turn; an interruption can stop the current response.

Read the connections in words
  • Spoken input → Agent: Understand and respond.
  • Agent → Speak: Deliver the response.
  • Speak → Listen again: New turn or interruption.
  • Listen again → Agent: feedback informs another turn.
Key idea

Conversation timing and interruption handling surround the model’s response.

CHOOSE YOUR PERSPECTIVE

Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Voice agents: see it in practice.

Agents that conduct spoken interactions while handling audio, timing, interruptions, and tool use.

What you’ll walk through

Follow a spoken interaction through interpretation, clarification, and a response or action. Focus on names, quantities, interruptions, and corrections that can change the user's intent.

The task in this version

Book a workshop place and confirm details first.

What you’ll learn to check

Transcript/audio controls, turn state, correction, scoped confirmation, and a clearly simulated booking result.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Everyday lifeAn authored case with its own evidence, changed condition, and decision.
The task in this example

Book a workshop place and confirm details first.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Transcript fixture: Saturday at ten. Slots: 10 am or 10 pm. Name heard as Lee or Leigh.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

This walkthrough represents speech as text. A real voice system must handle audio uncertainty and turn-taking as well as the task itself.

1 / 6

Apply this to your project

Describe your task to your own model and use Voice agents as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

A voice agent is a single agent you talk to instead of type to: speech in, speech out, in real time. Real time adds a decision chat does not need: something has to decide when the caller has stopped talking, and what happens if they start again while the agent is still going. OpenAI’s Realtime API documents voice activity detection that “will determine when the user has started or stopped speaking and respond automatically,” and for interruption: “The server will automatically truncate unplayed audio when there’s a user interruption”[1]. Google calls the same behavior “Barge-in”: “Users can interrupt the model at any time for responsive interactions”[2].

The example below is a text simulation, not a voice pipeline: nothing here records, streams, transcribes or synthesizes sound. What it can honestly show is the control flow (which choices are the model’s and which belong to the system around it) using Message content parts to reference audio without pretending to process it.

Level 5 still means the model decides the action and the stop: here, whether to keep talking or yield the floor. Your code enforces the caps: a latency budget on a turn, and an interruption that is never the model’s own choice to make.

This page is sourced, not measured: every latency and turn-taking claim below comes from a maker’s own documentation, and no call has been recorded and scored here. It is illustrated.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Voice agents

The model decides whether to keep talking; an interruption or a latency cap can cut it off first, and that part is never the model's choice.

Level 5 · Agent loops
Caller speaks (audio)Caller speaks(audio)MODELkeep talking, or yieldkeep talking,or yieldTOOLcontinue_speaking()continue_speaking()Floor returns to callerFloor returnsto caller
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 06Your code chose

The caller speaks

[audio] "What time do you close tonight?"
0 tokens · 0 ms

Practical guidance

If you are setting one up for your business rather than just talking to one, two things have to be in place before it takes a single real call: the disclosure, and the way out to a person.

Say plainly, before the caller can say anything else, that they are talking to an AI. ElevenLabs requires exactly this of businesses using its agents: telling callers “They are interacting with AI rather than a human” and that “Their conversations are being recorded and may be shared with ElevenLabs and its third-party large language model providers,”[4] with that notice “presented immediately prior to any interaction”[4], not after the first exchange. ElevenLabs is explicit that meeting this is the organization’s own responsibility, not legal advice[4], and what the law actually requires varies by place; check it for where you operate rather than assuming a platform’s minimum covers you.

Build the way out next: a phrase that reliably hands the call to a person (“talk to someone,” “I need a human”) and a real person or queue on the other end of it, not a dead end. Write the script for that handoff the same way you would write the disclosure: plainly, and tested before it goes live.

Then call it yourself, on your own phone, before a customer does. Interrupt it mid-sentence on purpose and check it actually stops: LiveKit’s own documentation describes a framework that “pauses the agent’s speech whenever it detects user speech in the input audio”[3], and a real deployment should behave the same way, not talk over you. Say a stray “mm-hmm” while it is mid-answer and check it keeps going rather than restarting, since a good one tells “true interruptions from conversational backchanneling” apart[3]. Ask for the handoff phrase and confirm it actually reaches your fallback, not a dropped call. Time the pause between finishing a sentence and its reply: a long silence with nothing said to fill it reads as broken, not thoughtful.

If what callers ask is small and known in advance (hours, an address, a balance) a script or a menu answers it without needing a conversation to manage at all.

Implementation details

The example simulates one turn’s control flow with a text model and no audio anywhere. The caller’s utterance is carried as an AudioPart label next to a TextPart transcript: the transcript is what the model actually reads, and the label is what a trace shows in place of sound it never processed, exactly the pattern examples/common/model.py documents for a run that never reaches a real speech model.

The model generates its answer in chunks. After each one it either calls continue_speaking, meaning it has more to say, or calls nothing, meaning it is done and the floor returns to the caller: decided_by: "model" either way, the same shape every level-5 example on this site uses for its stop. MAX_CHUNKS_PER_TURN (3) stands in for a latency budget: past a certain number of chunks a real system has to cut the agent off to stay responsive, whatever the model would have said next. interrupt_after_chunk simulates a caller starting to talk mid-turn; when it fires, the code cuts the agent off immediately, without asking the model anything: decided_by: "code", because a real interruption is a signal the system acts on the instant it arrives, not a choice the model gets to weigh in on. OpenAI’s own guidance for evaluating a voice agent puts this on the list of things to measure: “audible response timing, unwanted silence, overlap, and yielding to interruptions”[5].

examples/voice_agents/run.py · lines 38–86
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    interrupt_after_chunk: int | None = None,
    max_chunks: int = MAX_CHUNKS_PER_TURN,
    max_tokens: int = MAX_TOKENS,
) -> Answer:
    del embedder  # no retrieval here; this page is about turn control, not what gets said
    content = [AudioPart(media_type="audio/wav", label=question), TextPart(text=question)]
    tracer.record(kind="code", decided_by="code", title="Caller speaks", detail=f"[audio] {question[:150]}")
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=content)]

    chunks: list[str] = []
    tokens_used = 0
    for chunk_index in range(max_chunks):
        if interrupt_after_chunk is not None and chunk_index == interrupt_after_chunk:
            tracer.record(
                kind="code", decided_by="code", title="Caller interrupts; code cuts the agent off",
                detail=f"stopped after {len(chunks)} chunk(s)",
            )
            break

        completion = model.complete(messages, tools=[CONTINUE_TOOL], max_tokens=100)
        tokens_used += completion.tokens_in + completion.tokens_out
        chunks.append(completion.text)

        if not completion.tool_calls:
            record_completion(tracer, decided_by="model", title="Model finishes and yields the floor", completion=completion)
            break

        record_completion(tracer, decided_by="model", title="Model chooses to keep talking", completion=completion)
        messages.append(Message(role="assistant", content=completion.text))

        if tokens_used >= max_tokens:
            tracer.record(
                kind="code", decided_by="code", title="Latency budget reached; code cuts the agent off",
                detail=f"{tokens_used} >= {max_tokens} tokens",
            )
            break
    else:
        tracer.record(
            kind="code", decided_by="code", title="Chunk cap reached; code cuts the agent off",
            detail=f"{max_chunks} chunks",
        )

    return Answer(text=" ".join(c for c in chunks if c), citations=[])

The same control flow answers a hands-busy question at a bench: an engineer with both hands full asks for the last reading and expects to hear it back, not a new one. That is honest only if the voice layer never produces the number itself. SYSTEM_PROMPT here is about turn control, not what gets said, so a bench version would have to add one more rule: read back a value the caller is given, never estimate, recall or restate one from anywhere else.

Run it yourself:

examples/voice_agents/README.md · lines 18–18
python -m examples.voice_agents --model stub:scripted
When you do not need this

Try plain chat first if the interaction does not actually need to happen in real time: turn-taking, interruption and latency budgets are all cost you pay for synchrony you may not need.

Try a fixed workflow instead (transcribe the utterance, then route it to one of a few known intents, then answer with a scripted response) if the things a caller might ask are small and known in advance; that is cheaper and does not need the model to manage the turn itself.

Move up to a voice agent once the range of things a caller might say is too open for a fixed set of intents, and the conversation genuinely needs to flow rather than follow a menu.

Failure modes

The agent talks over the caller

How to notice it
The caller starts speaking and the agent keeps going instead of yielding immediately, breaking the sense that anyone is actually listening.
How to test for it
Interrupt mid-sentence and time how long the agent keeps talking before it stops. LiveKit documents its framework pausing agent speech the instant it detects caller speech; a noticeable delay past that is this failure.

A false interruption derails the agent

How to notice it
A stray "mm-hmm" or a cough gets read as a real interruption, and the agent restarts or drops what it was saying instead of continuing.
How to test for it
Say a short backchannel sound while the agent is mid-answer and check whether it resumes from where it left off, which is the recovery LiveKit's documentation describes, or restarts from scratch.

No disclosure, or disclosure too late

How to notice it
The caller is well into the conversation before anything tells them they are talking to an AI, if anything ever does.
How to test for it
Start a session and check whether a disclosure plays before you can say anything at all. ElevenLabs requires this to be presented immediately prior to any interaction, not after the first exchange.

The latency budget is blown silently

How to notice it
A turn takes long enough that the pause reads as dead air, with nothing telling the caller the agent is still working.
How to test for it
Measure the gap between the caller finishing and the agent's first audible response across many turns, not just once; an occasional slow turn with no filler or acknowledgment is this failure even if the average looks fine.

The cap cuts off mid-sentence with no recovery

How to notice it
A chunk or token cap ends the turn partway through a sentence, and the agent neither finishes the thought nor says anything to cover the abrupt stop.
How to test for it
Force a low chunk cap on an answer that needs more than one chunk (this page's own test suite does exactly this) and check whether what comes back reads as a real, if short, answer or as speech cut off mid-word.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1Model calls, one-chunk turn
3Model calls, full chunk cap reached
~180–260Tokens in, one chunk
~0.4sWall time, one chunk (illustrated)
Compared with a text chat turn (level 1)A real voice turn adds speech-to-text and text-to-speech latency on top of whatever this diagram shows, neither of which this text-only example measures; the illustrated numbers cover only the control-flow decisions themselves.

How to Evaluate It

voice_agents does not do the site’s own question-answering task and is not part of the shared 60-question set: a text-simulated turn loop has no document to cite. What changes for evaluation here is the task itself: voice needs its own small eval, on its own page, the way this site treats every technique whose running task is not a fair test of it. What that eval would measure is specific to a spoken conversation: yielding to interruptions (does the agent actually stop when talked over), time from the caller finishing to the agent’s first audible response, and false-interrupt recovery (does a backchannel derail it). scripts/eval_run.py knows voice_agents and refuses to score it, printing that reason; no runner for the measures above exists yet.

Run it

What to monitor

How quickly the agent actually stops when talked over, the false-interrupt rate, and response latency per turn, tracked separately from whatever the underlying model's own latency is.

Cost at volume

A real deployment pays for speech-to-text and text-to-speech on top of every model call this example's control flow makes, often the larger share of per-minute cost at volume, not the model itself.

How it fails in production

A change to the underlying model shifts its per-chunk timing enough that the fixed latency budget starts cutting off answers that used to finish in time, with nothing in a dashboard built around token counts likely to show it.

What to log

Every chunk generated, the time between them, whether a turn ended by the model yielding, an interruption, or a cap, and whether whatever disclosure your platform or your lawyers require actually played before the conversation started.

Try it

  1. Use it

    Start a conversation with a voice assistant and interrupt it mid-sentence on purpose. Does it stop immediately? Does anything at the start of the call tell you it is an AI?

  2. Build it

    Run python -m examples.voice_agents --model stub:scripted from the repo root. The model speaks one chunk, keeps the floor, speaks a second, and yields on its own: two decisions, not one reply. Now give every entry in SCRIPTED (examples/voice_agents/__main__.py) a continue_speaking call and add a third. Nothing yields, so the code cuts the agent off at the cap.

  3. Either lane

    Write the disclosure sentence you would want at the start of a call with an AI. Then check the SYSTEM_PROMPT in examples/voice_agents/run.py: it does nothing like it, on purpose, since this example is about turn control. Where would you add it?

How it connects

Before, after and instead of this

Optional: products, tools, and models

8 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 2 more examples
  • Pipecat Tool or framework · Daily

    Voice agent framework

    Checked 09/18/2026
  • Vapi Tool or framework · Vapi

    Voice agent platform

    Checked 09/18/2026
In practice

Handle a spoken support request

The agent listens, responds aloud, and yields when the caller interrupts to correct a detail.

Out there

Named products, tools and models

Products3
  • ChatGPT voiceOpenAI · voice assistant
  • ElevenLabsElevenLabs · voice generation
  • Gemini LiveGoogle · voice assistant
Tools5
  • Gemini Live APIGoogle · voice API
  • LiveKit AgentsLiveKit · voice agent framework
  • OpenAI Realtime APIOpenAI · voice API
  • PipecatDaily · voice agent framework
  • VapiVapi · voice agent platform

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Realtime conversations · OpenAI (API documentation) (accessed 09/19/2026)
  2. Gemini Live API overview (archived copy) · Google (Gemini API documentation, via the Internet Archive) (accessed 09/19/2026)
  3. Turns overview · LiveKit (accessed 09/19/2026)
  4. Disclosure requirements · ElevenLabs (accessed 09/19/2026)
  5. Voice agents · OpenAI (API documentation) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page