The example simulates one turn’s control flow with a text model and no audio anywhere. The caller’s
utterance is carried as an AudioPart label next to a TextPart transcript: the transcript is
what the model actually reads, and the label is what a trace shows in place of sound it never
processed, exactly the pattern examples/common/model.py documents for a run that never reaches a
real speech model.
The model generates its answer in chunks. After each one it either calls continue_speaking,
meaning it has more to say, or calls nothing, meaning it is done and the floor returns to the
caller: decided_by: "model" either way, the same shape every level-5 example on this site uses
for its stop. MAX_CHUNKS_PER_TURN (3) stands in for a latency budget: past a certain number of
chunks a real system has to cut the agent off to stay responsive, whatever the model would have
said next. interrupt_after_chunk simulates a caller starting to talk mid-turn; when it fires, the
code cuts the agent off immediately, without asking the model anything: decided_by: "code",
because a real interruption is a signal the system acts on the instant it arrives, not a choice the
model gets to weigh in on. OpenAI’s own guidance for evaluating a voice agent puts this on the
list of things to measure: “audible response timing, unwanted silence, overlap, and yielding to
interruptions”[5].
examples/voice_agents/run.py · lines 38–86
def run(
question: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
interrupt_after_chunk: int | None = None,
max_chunks: int = MAX_CHUNKS_PER_TURN,
max_tokens: int = MAX_TOKENS,
) -> Answer:
del embedder # no retrieval here; this page is about turn control, not what gets said
content = [AudioPart(media_type="audio/wav", label=question), TextPart(text=question)]
tracer.record(kind="code", decided_by="code", title="Caller speaks", detail=f"[audio] {question[:150]}")
messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=content)]
chunks: list[str] = []
tokens_used = 0
for chunk_index in range(max_chunks):
if interrupt_after_chunk is not None and chunk_index == interrupt_after_chunk:
tracer.record(
kind="code", decided_by="code", title="Caller interrupts; code cuts the agent off",
detail=f"stopped after {len(chunks)} chunk(s)",
)
break
completion = model.complete(messages, tools=[CONTINUE_TOOL], max_tokens=100)
tokens_used += completion.tokens_in + completion.tokens_out
chunks.append(completion.text)
if not completion.tool_calls:
record_completion(tracer, decided_by="model", title="Model finishes and yields the floor", completion=completion)
break
record_completion(tracer, decided_by="model", title="Model chooses to keep talking", completion=completion)
messages.append(Message(role="assistant", content=completion.text))
if tokens_used >= max_tokens:
tracer.record(
kind="code", decided_by="code", title="Latency budget reached; code cuts the agent off",
detail=f"{tokens_used} >= {max_tokens} tokens",
)
break
else:
tracer.record(
kind="code", decided_by="code", title="Chunk cap reached; code cuts the agent off",
detail=f"{max_chunks} chunks",
)
return Answer(text=" ".join(c for c in chunks if c), citations=[])
The same control flow answers a hands-busy question at a bench: an engineer with both hands full
asks for the last reading and expects to hear it back, not a new one. That is honest only if the
voice layer never produces the number itself. SYSTEM_PROMPT here is about turn control, not
what gets said, so a bench version would have to add one more rule: read back a value the caller
is given, never estimate, recall or restate one from anywhere else.
Run it yourself:
examples/voice_agents/README.md · lines 18–18
python -m examples.voice_agents --model stub:scripted