Voice AI & Conversational AgentsVoice agent foundations · Lesson 2 of 17
Latency budgets, turn-taking and barge-in
Video lecture
Latency budgets, turn-taking and barge-in
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Latency and turn-taking
Here's a number worth remembering: in everyday conversation across many languages, researchers found the typical gap between one person finishing and the other starting is only around a fifth of a second. We're wired for fast turn-taking. When a voice agent leaves a two-second silence, callers repeat themselves, talk over it, or hang up. In this lesson, you'll learn to budget latency, cut it, and handle the hardest part of voice: knowing when the caller has actually finished, and what to do when they interrupt.
0:37 Why timing matters
Why does latency deserve a whole lesson? Because on the phone, timing is the experience. Callers forgive a slightly imperfect answer delivered naturally, but they abandon a perfect answer delivered after an awkward silence. Slow agents also cause more interruptions, which confuse turn detection and create a spiral of talking over each other. Fixing latency and turn-taking often improves satisfaction more than any prompt change.
1:05 The latency budget
Voice to voice latency is the time from the caller finishing to hearing the first sound of the reply. In a cascade, it's the sum of end of turn detection, final speech recognition, the language model's time to first token, plus any tool call, synthesis time to first audio, and network and telephony transport. An illustrative budget for a phone agent lands somewhere between about point seven and two seconds. Your numbers will differ, so measure them. And tool calls add their own time. A slow CRM lookup can add whole seconds.
1:45 Analogy: merging onto a motorway
An analogy: turn-taking is like merging onto a busy motorway. Merge too early and you cut someone off. That's the agent interrupting a caller who paused to think. Wait too long and the cars behind you start honking. That's dead air after a crisp answer. Good drivers read signals: indicators, speed, the gap. Good turn detection reads signals too: silence, whether the sentence sounds finished, and what you just asked for. If you asked for a phone number, you expect pauses between digit groups.
2:22 Six latency levers
Six levers cut latency. Stream everything: partial transcripts, token streaming, and synthesis that starts on the first phrase. Pick fast components for conversation, like low-latency voice models and fast language model tiers, in regions near your callers. Keep prompts and retrieved context lean, because long context slows the first token. Use speculative work, starting a likely reply before the turn is confirmed. Make tools fast, with caching and tight timeouts. And put your tool servers in the same region as the agent platform.
2:58 When is the caller done?
Now the hardest problem: when has the caller finished? A fixed silence timer fails both ways. It cuts off people who pause to think, like someone reading out an account number, and feels sluggish after a crisp yes. Modern systems combine voice activity detection, which asks, is there speech right now, with semantic end of turn detection, which asks, does this sound complete? OpenAI's Realtime API has a semantic V A D mode. ElevenAgents has settings like turn eagerness and patience for spelled-out input. And context helps: if you just asked for a phone number, expect digits with pauses.
3:41 Barge-in
Then there's barge-in. Callers interrupt: no, no, Thursday, not Tuesday. A good agent stops quickly, drops the rest of its sentence, and listens. Allow interruptions by default, though you might briefly disable them for a legally required statement, then invite the caller to speak. Ignore backchannels: mm-hmm or okay while the agent talks usually means carry on, not stop. And watch for echo on speakerphones, where the agent hears itself and interrupts itself. Always test on real phones.
4:15 Worked example: Lahore insurance agent
Here's a tuning story from an insurance callback agent in Lahore. In the pilot, callers reading policy numbers got cut off, the agent felt slow after yes or no questions, and database lookups caused three-second silences. The team increased patience for numeric input and had the agent confirm digits in groups. They made turn-taking more eager on confirmation steps. They added a short line before lookups, cached policy summaries, and moved their API into the same region. And they added Urdu and English backchannel words, like haan, acha and okay, to the ignore list.
4:56 Measure it
How do you know it's working? Measure. Most platforms expose per-turn timing. Track median and ninetieth percentile voice to voice latency, separately for turns with and without tools, because averages hide the painful long pauses. Track how often callers talk over the agent, which signals slowness or verbosity. Track false barge-ins, where the agent stops for a backchannel or echo. And track cut-offs, where the agent jumps in mid-thought. The lesson text has a small Python script that computes these from exported turn data.
5:33 Example 2: Dubai pizza orders
A simple example. A pizza shop in Dubai tests its ordering agent and finds it always replies about two seconds after a caller says yes. The cause: the agent's system prompt was over two thousand words, and it retrieved five long menu chunks on every turn. The fix: trim the prompt to the essentials, retrieve only the one relevant menu section, and switch to a faster model tier for confirmation turns. The reply after yes now feels instant, and nothing about the agent's knowledge was lost.
6:10 Common mistakes
Common mistakes with latency. Measuring averages instead of the ninetieth percentile, which hides the long pauses callers remember. Adding slow tools without a filler line, so callers hear silence. Tuning turn detection only in a quiet office on a laptop, when real callers are in cars, markets and on speakerphones. And fixing one number, like model speed, while ignoring the tool call that's actually causing most of the delay.
6:40 Watch me do it: tune turn-taking
Watch me do it. I open the agent config and fill the turn taking section. First, the baseline: I run ten test calls and export per-turn timings, then run the latency script. The median looks fine, but the ninetieth percentile is much higher, and every slow turn involves the availability tool. So I make three changes. One: in the tool definition I switch on pre-tool speech, so the agent says let me check that for you while waiting. Two: I set turn eagerness to normal for most of the call, but I add a note in the prompt that when collecting phone numbers or dates, the agent should wait for the caller to finish and confirm digits in groups. The platform also has a patience setting for spelled input, which I raise. Three: I add backchannel words to the ignore list: okay, mm-hmm, yeah, and for our Arabic callers, aiwa and tamam. Then I run the same ten calls again and re-run the script. I record the new median and ninetieth percentile next to the old ones in the config comments. The tool turns still take longest, but callers now hear a natural phrase instead of silence, and nobody was cut off mid-number. Next lesson, we'll choose the voice itself.
8:11 Recap and next step
Recap. Callers expect near-human timing. Budget latency across detection, recognition, the model, tools, synthesis and transport. Cut it by streaming, choosing fast components, keeping context lean and making tools fast. Use semantic turn detection and context-aware patience, allow interruptions but ignore backchannels, and measure median and p90. Your next step: export turn timings from ten test calls and run the script. Find your slowest stage and pick one lever to pull.
Why milliseconds matter
In human conversation, the gap between one person finishing and the other starting is famously short: cross-language research (Stivers and colleagues, 2009) found typical gaps of around a fifth of a second. We start planning our reply before the other person finishes. Voice agents cannot yet match that consistently, but callers notice when gaps stretch toward two seconds: they repeat themselves, talk over the agent, or hang up.
The latency budget
"Voice-to-voice latency" is the time from the caller finishing speaking to hearing the first audio of the reply. In a cascade it is roughly:
voice-to-voice = end-of-turn detection
+ final ASR
+ LLM time-to-first-token (+ tool call time, if any)
+ TTS time-to-first-audio
+ network and telephony transport (both directions)An illustrative budget for a cascaded phone agent (your numbers will differ; measure them):
| Stage | Illustrative target |
|---|---|
| End-of-turn detection | 200 to 500 ms |
| Final ASR after end of speech | 100 to 300 ms |
| LLM time to first token | 200 to 700 ms |
| TTS time to first audio | 75 to 300 ms |
| Network and telephony | 100 to 300 ms |
| Total | roughly 0.7 to 2 seconds |
Tool calls add their own time: a slow CRM lookup can add seconds. That is why agents speak a short line before slow tools.
Levers to cut latency
- Stream everything: partial ASR, token streaming from the LLM, TTS that starts on the first phrase.
- Pick fast components for conversation: low-latency TTS models (for example ElevenLabs' Flash v2.5), fast LLM tiers, regional endpoints near your callers and telephony.
- Keep prompts and context lean: long system prompts and huge retrieved chunks slow time-to-first-token.
- Speculative work: start generating a likely reply before the turn is fully confirmed, and discard it if the caller keeps talking (some platforms offer a "speculative turn" setting).
- Fast tools: cache availability, precompute common lookups, set tight timeouts with graceful fallbacks.
- Co-locate: run your tool servers in the same region as the agent platform.
Turn-taking: when has the caller finished?
The naive approach waits for a fixed silence (say 700 ms). That fails in both directions: it cuts off people who pause to think ("My account number is... 4 4 7...") and feels sluggish for crisp answers ("Yes.").
Modern systems combine:
- VAD: is there speech energy now?
- Semantic end-of-turn detection: does what they said sound complete? OpenAI's Realtime API offers a
semantic_vadturn-detection mode that waits longer when speech trails off and responds faster to definitive statements. ElevenAgents exposes turn settings such as turn eagerness (patient, normal, eager) and patience for spelled-out input. LiveKit Agents and Pipecat ship open turn-detection models. - Context: when the agent just asked for a phone number, expect digits with pauses.
Barge-in (interruptions)
Callers interrupt: "No, no, Thursday, not Tuesday." A good agent stops speaking quickly, discards the unspoken remainder, and listens. Design choices:
- Allow interruptions by default, but consider disabling them briefly for legally required statements (for example a recording notice) and re-offering the chance to speak afterwards.
- Ignore backchannels: "mm-hmm", "okay", "yeah" while the agent talks usually mean "go on", not "stop". Platforms increasingly filter these (for example configurable ignore terms).
- Echo cancellation: on speakerphones, the agent can hear itself and "interrupt" itself. Telephony and WebRTC stacks include echo cancellation; test on real devices.
Worked example: tuning an insurance callback agent in Lahore
Symptoms in the pilot: callers reading policy numbers were cut off mid-number; the agent felt slow after "yes/no" questions; tool lookups caused 3-second silences.
Fixes:
| Problem | Fix |
|---|---|
| Cut off during policy numbers | Increased patience for spelled or numeric input; prompt tells agent to confirm digits in groups |
| Slow after short answers | Switched turn eagerness to a more eager setting for confirmation steps |
| 3-second silences on lookups | Added pre-tool speech ("One moment while I pull that up"), cached policy summaries, moved API to the same region |
| Agent stopped when caller said "hmm" | Added backchannel ignore terms in Urdu and English ("haan", "acha", "okay") |
Median voice-to-voice latency dropped, and abandoned calls fell noticeably in the next week's sample. (Measure your own before-and-after rather than trusting generic claims.)
Hands-on: measure latency from call logs
Most platforms expose per-turn timing metrics in conversation details or via API. A simple analysis:
import statistics, json
# turns.json: [{"turn": 1, "user_end_ms": 10250, "agent_audio_start_ms": 11320, "tool_ms": 0}, ...]
turns = json.load(open("turns.json"))
v2v = [t["agent_audio_start_ms"] - t["user_end_ms"] for t in turns]
with_tools = [t["agent_audio_start_ms"] - t["user_end_ms"] for t in turns if t["tool_ms"] > 0]
def pct(values, p):
values = sorted(values)
k = max(0, min(len(values) - 1, round(p / 100 * (len(values) - 1))))
return values[k]
print("median v2v ms:", statistics.median(v2v))
print("p90 v2v ms:", pct(v2v, 90))
if with_tools:
print("median v2v with tools ms:", statistics.median(with_tools))Track the median and 90th percentile; averages hide the painful long pauses.
Measuring success
- Median and p90 voice-to-voice latency, separately for turns with and without tools.
- Interruption rate: how often callers talk over the agent (high rates signal slowness or verbosity).
- False barge-in rate: agent stops for backchannels or echo.
- Cut-off rate: agent starts speaking while the caller was mid-thought.
Key takeaways
- Voice-to-voice latency sums turn detection, ASR, LLM first token, tools, TTS first audio and transport; measure median and p90.
- Stream every stage, choose low-latency components and regions, keep context lean, and make tools fast with pre-tool speech.
- Combine VAD with semantic end-of-turn detection and context-aware patience for numbers and spelling.
- Allow barge-in, ignore backchannels, handle echo, and track talk-over, false barge-in and cut-off rates.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Run ten test calls, export per-turn timings, and compute median and p90 voice-to-voice latency with and without tools. Pick one lever to improve the slowest stage.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.