Voice AI & Conversational AgentsVoice agent foundations · Lesson 2 of 17

Latency budgets, turn-taking and barge-in

Article · 15 min · 9 min lecture

Video lecture

Latency budgets, turn-taking and barge-in

13 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 13

Latency and turn-taking

  • Human turn gaps: about a fifth of a second
  • Long silences lose callers
  • Budget, cut, detect, handle interruptions

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why milliseconds matter

In human conversation, the gap between one person finishing and the other starting is famously short: cross-language research (Stivers and colleagues, 2009) found typical gaps of around a fifth of a second. We start planning our reply before the other person finishes. Voice agents cannot yet match that consistently, but callers notice when gaps stretch toward two seconds: they repeat themselves, talk over the agent, or hang up.

The latency budget

"Voice-to-voice latency" is the time from the caller finishing speaking to hearing the first audio of the reply. In a cascade it is roughly:

voice-to-voice = end-of-turn detection
               + final ASR
               + LLM time-to-first-token (+ tool call time, if any)
               + TTS time-to-first-audio
               + network and telephony transport (both directions)

An illustrative budget for a cascaded phone agent (your numbers will differ; measure them):

StageIllustrative target
End-of-turn detection200 to 500 ms
Final ASR after end of speech100 to 300 ms
LLM time to first token200 to 700 ms
TTS time to first audio75 to 300 ms
Network and telephony100 to 300 ms
Totalroughly 0.7 to 2 seconds

Tool calls add their own time: a slow CRM lookup can add seconds. That is why agents speak a short line before slow tools.

Levers to cut latency

  1. Stream everything: partial ASR, token streaming from the LLM, TTS that starts on the first phrase.
  2. Pick fast components for conversation: low-latency TTS models (for example ElevenLabs' Flash v2.5), fast LLM tiers, regional endpoints near your callers and telephony.
  3. Keep prompts and context lean: long system prompts and huge retrieved chunks slow time-to-first-token.
  4. Speculative work: start generating a likely reply before the turn is fully confirmed, and discard it if the caller keeps talking (some platforms offer a "speculative turn" setting).
  5. Fast tools: cache availability, precompute common lookups, set tight timeouts with graceful fallbacks.
  6. Co-locate: run your tool servers in the same region as the agent platform.

Turn-taking: when has the caller finished?

The naive approach waits for a fixed silence (say 700 ms). That fails in both directions: it cuts off people who pause to think ("My account number is... 4 4 7...") and feels sluggish for crisp answers ("Yes.").

Modern systems combine:

  • VAD: is there speech energy now?
  • Semantic end-of-turn detection: does what they said sound complete? OpenAI's Realtime API offers a semantic_vad turn-detection mode that waits longer when speech trails off and responds faster to definitive statements. ElevenAgents exposes turn settings such as turn eagerness (patient, normal, eager) and patience for spelled-out input. LiveKit Agents and Pipecat ship open turn-detection models.
  • Context: when the agent just asked for a phone number, expect digits with pauses.

Barge-in (interruptions)

Callers interrupt: "No, no, Thursday, not Tuesday." A good agent stops speaking quickly, discards the unspoken remainder, and listens. Design choices:

  • Allow interruptions by default, but consider disabling them briefly for legally required statements (for example a recording notice) and re-offering the chance to speak afterwards.
  • Ignore backchannels: "mm-hmm", "okay", "yeah" while the agent talks usually mean "go on", not "stop". Platforms increasingly filter these (for example configurable ignore terms).
  • Echo cancellation: on speakerphones, the agent can hear itself and "interrupt" itself. Telephony and WebRTC stacks include echo cancellation; test on real devices.

Worked example: tuning an insurance callback agent in Lahore

Symptoms in the pilot: callers reading policy numbers were cut off mid-number; the agent felt slow after "yes/no" questions; tool lookups caused 3-second silences.

Fixes:

ProblemFix
Cut off during policy numbersIncreased patience for spelled or numeric input; prompt tells agent to confirm digits in groups
Slow after short answersSwitched turn eagerness to a more eager setting for confirmation steps
3-second silences on lookupsAdded pre-tool speech ("One moment while I pull that up"), cached policy summaries, moved API to the same region
Agent stopped when caller said "hmm"Added backchannel ignore terms in Urdu and English ("haan", "acha", "okay")

Median voice-to-voice latency dropped, and abandoned calls fell noticeably in the next week's sample. (Measure your own before-and-after rather than trusting generic claims.)

Hands-on: measure latency from call logs

Most platforms expose per-turn timing metrics in conversation details or via API. A simple analysis:

import statistics, json

# turns.json: [{"turn": 1, "user_end_ms": 10250, "agent_audio_start_ms": 11320, "tool_ms": 0}, ...]
turns = json.load(open("turns.json"))
v2v = [t["agent_audio_start_ms"] - t["user_end_ms"] for t in turns]
with_tools = [t["agent_audio_start_ms"] - t["user_end_ms"] for t in turns if t["tool_ms"] > 0]

def pct(values, p):
    values = sorted(values)
    k = max(0, min(len(values) - 1, round(p / 100 * (len(values) - 1))))
    return values[k]

print("median v2v ms:", statistics.median(v2v))
print("p90 v2v ms:", pct(v2v, 90))
if with_tools:
    print("median v2v with tools ms:", statistics.median(with_tools))

Track the median and 90th percentile; averages hide the painful long pauses.

Measuring success

  • Median and p90 voice-to-voice latency, separately for turns with and without tools.
  • Interruption rate: how often callers talk over the agent (high rates signal slowness or verbosity).
  • False barge-in rate: agent stops for backchannels or echo.
  • Cut-off rate: agent starts speaking while the caller was mid-thought.

Key takeaways

  • Voice-to-voice latency sums turn detection, ASR, LLM first token, tools, TTS first audio and transport; measure median and p90.
  • Stream every stage, choose low-latency components and regions, keep context lean, and make tools fast with pre-tool speech.
  • Combine VAD with semantic end-of-turn detection and context-aware patience for numbers and spelling.
  • Allow barge-in, ignore backchannels, handle echo, and track talk-over, false barge-in and cut-off rates.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Callers are cut off while reading account numbers. What is the best first fix?
  2. Why track p90 latency rather than only the average?
  3. While the agent speaks, the caller says 'mm-hmm'. What should usually happen?

Put it into practice

Run ten test calls, export per-turn timings, and compute median and p90 voice-to-voice latency with and without tools. Pick one lever to improve the slowest stage.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.