Multimodal & Reasoning Models in PracticeAudio: speech-to-text and text-to-speech · Lesson 7 of 17
Realtime voice agents: speech-to-speech in practice
Video lecture
Realtime voice agents: speech-to-speech in practice
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Realtime voice agents
Call a dental clinic in Dubai at nine at night, and an assistant answers in your language, finds a slot for tomorrow, confirms the dentist, and books it, all in under a minute, and it lets you interrupt when you change your mind. That is a realtime voice agent. In this lecture you will learn the two architectures for voice agents, the building blocks every one needs, how to prompt for the ear, how a realtime session is configured, how to budget latency, and the compliance rules for voice.
0:39 Two architectures
Voice agents used to be built as a chain: speech-to-text, then a language model, then text-to-speech. Each stage added delay, and tone of voice was lost in the middle. Realtime speech-to-speech models now take audio in and produce audio out directly, over a persistent connection, typically a WebSocket or WebRTC. Examples include OpenAI's Realtime API and Google's Gemini Live API. The component pipeline is still widely used when you need a particular speech vendor, maximum control, specific languages, or text logs at every stage.
1:16 Trade-offs
Compare them honestly. A pipeline lets you choose best-in-class parts, log text at each stage, and pick from many voices and languages, but you must manage more latency, and emotion is lost between stages. A realtime model gives lower latency, natural turn-taking and preserved tone, and handles interruptions well, but offers fewer voice choices, less visibility into intermediate text, and more dependence on one provider. Many teams prototype with a realtime model and keep a pipeline for languages or voices it does not cover.
1:53 Four building blocks
Every voice agent needs four building blocks. Turn detection, also called voice activity detection, decides when the user has finished speaking. Too eager and the agent interrupts; too slow and it feels laggy. Most realtime APIs offer server-side detection with tunable sensitivity. Barge-in lets the user interrupt, and the agent must stop speaking and listen immediately. Tool calls mid-conversation, like checking a booking system, should be covered by a short filler phrase. And transport: browsers usually connect with WebRTC, and phone lines through SIP or a telephony provider.
2:31 Prompting for the ear
Voice prompts need everything a text prompt needs, plus rules for how things sound. For the dental clinic: reply in the caller's language, English or Arabic. Warm, calm and concise, one or two short sentences per turn. Never read out lists longer than three items; offer to send details by text message. Say times naturally, like ten thirty in the morning. Before booking, confirm the date, time and dentist in one sentence and wait for a clear yes. Boundaries: say you are an automated assistant at the start of every call; give no medical advice; for pain, swelling or emergencies, give the emergency number and offer a transfer; and transfer if the caller asks for a person twice.
3:22 Configuring a session
How is a realtime session configured? After opening the connection, you send a session update event with the model, the instructions, the audio settings such as server-side turn detection and the output voice, and your tools. In the lesson, the tool is find slots, which takes a dentist and a date in year-month-day format. Field names evolve, so confirm against the current reference. Most teams start from the provider's official realtime examples or a voice-agent framework, open-source or hosted platforms such as ElevenLabs Agents, which handle transport, telephony and turn-taking, then add their own prompt, tools and logging.
4:05 Latency budget
Latency makes or breaks voice. Measure each segment: the moment the user stops speaking, turn detected, first model audio, and each tool round trip. Keep a latency budget per turn and test on mobile networks, not just office Wi-Fi. The biggest perceived wins often come from short first responses and parallel tool calls, rather than from swapping models. And log everything with timestamps, so when a caller complains about a long pause, you can see exactly which segment caused it.
4:40 Compliance and measurement
Compliance matters from the first word. Disclose automation at the start of calls. Record and transcribe only with appropriate notice or consent under local law, and keep retention short. For outbound calls, check telemarketing and AI-voice rules in each market, such as US rules on AI-generated voices in robocalls and consumer protection rules in the UK, the UAE and Saudi Arabia. And always offer a route to a human. Then measure: task completion, turn latency, interruption handling, transfer rate and caller satisfaction. Review a sample of calls weekly with the prompt open beside you; most improvements come from small prompt and tool fixes found that way.
5:26 Example 1: a restaurant booking
A simple worked example. A small restaurant wants a voice assistant that answers one question: can I book a table for tonight? First version: the agent reads out every available slot, all eleven of them, in one long sentence. Callers hang up. Second version, with voice-first rules: offer at most two options, the closest to what the caller asked; confirm the time and number of people in one sentence; wait for a clear yes before booking. Now the call goes: for four people tonight I have seven or eight thirty; which suits you? Short, natural, done in under a minute.
6:09 Example 2: physio clinic bookings (illustrative)
Now a business scenario, with illustrative numbers. A chain of physiotherapy clinics in the UK receives about nine thousand calls a month, a third outside opening hours. They pilot a realtime voice agent for booking, rescheduling and cancelling. In the first week, median response latency after a caller stops speaking is about one point eight seconds, and callers talk over the agent. The team tunes turn detection, adds a short filler, let me check that, while the booking tool runs, and runs tool calls in parallel. Median latency falls to about zero point nine seconds. They also add a transfer to reception after two failed attempts, and a disclosure at the start of every call. After a month, about sixty percent of out-of-hours booking calls complete without staff, and satisfaction scores are similar to human-handled calls. Illustrative figures.
7:09 Recap
To recap. Choose between a component pipeline and a realtime speech-to-speech model based on control, voices, languages and latency. Get turn detection, barge-in, filler phrases and transport right. Write voice prompts with style, task and boundary sections, confirm before actions, and disclose automation. Measure latency per segment and review real calls. Try this now: write a voice-agent system prompt for a real booking or support use case, test five calls, and log the latency of each segment. Next module: image and video generation.
From pipelines to speech-to-speech
Voice agents used to be built as a chain: speech-to-text, then a language model, then text-to-speech. Each stage added delay, and tone of voice was lost in the middle. Realtime speech-to-speech models now take audio in and produce audio out directly, over a persistent connection (typically WebSocket or WebRTC). Examples include OpenAI's Realtime API with its realtime voice models and Google's Gemini Live API. The component pipeline is still widely used, especially when you need a particular STT or TTS vendor, maximum control or specific languages.
| Approach | Strengths | Trade-offs |
|---|---|---|
| Component pipeline (STT, LLM, TTS) | Choose best-in-class parts, easy to log text at each stage, flexible languages and voices | More latency to manage, tone and emotion lost between stages |
| Speech-to-speech realtime model | Lower latency, natural turn-taking, preserves tone, handles interruptions | Fewer voice choices, less visibility into intermediate text, provider lock-in |
The building blocks of any voice agent
- Turn detection (voice activity detection): deciding when the user has finished speaking. Too eager and the agent interrupts; too slow and it feels laggy. Most realtime APIs offer server-side turn detection with tunable sensitivity.
- Barge-in: letting the user interrupt the agent, which must stop speaking immediately and listen.
- Tool calls mid-conversation: checking a booking system, then speaking the answer. Say a short filler ("Let me check that") while tools run.
- Telephony and transport: browsers via WebRTC, phone lines via SIP or a telephony provider.
Prompting a voice agent
Voice prompts need everything a text system prompt needs, plus rules for how things sound:
You are the booking assistant for Harbour Dental Clinic in Dubai.
You speak on the phone with patients in English or Arabic; reply in the
language the caller uses.
Style:
- Warm, calm and concise. One or two short sentences per turn.
- Never read out lists longer than three items; offer to send details by SMS.
- Say times naturally ("ten thirty in the morning"), not "10:30".
- Confirm names by spelling back only if the caller spells them first.
Task:
- Help callers book, move or cancel appointments using the tools.
- Before booking, confirm the date, time and dentist in one sentence and
wait for a clear "yes".
Boundaries:
- You are an automated assistant; say so at the start of every call.
- Do not give medical advice. For pain, swelling or emergencies, give the
emergency number and offer to transfer to reception.
- If the caller is frustrated or asks for a person twice, transfer.Hands-on: a realtime session configuration
Realtime APIs are configured at session start with instructions, voice, turn detection and tools. The shape below follows the OpenAI Realtime API's session update event; field names evolve, so confirm against the current reference before use.
import json, os
# Sent as the first event after opening the realtime WebSocket or WebRTC data channel
session_update = {
"type": "session.update",
"session": {
"type": "realtime",
"model": os.environ["REALTIME_MODEL"],
"instructions": open("prompts/harbour_dental_voice.md", encoding="utf-8").read(),
"audio": {
"input": {"turn_detection": {"type": "server_vad"}},
"output": {"voice": os.environ.get("REALTIME_VOICE", "marin")},
},
"tools": [{
"type": "function", "name": "find_slots",
"description": "Find free appointment slots for a dentist on a date (YYYY-MM-DD).",
"parameters": {"type": "object", "properties": {
"dentist": {"type": "string"}, "date": {"type": "string"}},
"required": ["dentist", "date"]},
}],
},
}
payload = json.dumps(session_update) # send over the open connectionMost teams start from the provider's official realtime examples or a voice-agent framework (open-source frameworks and hosted voice-agent platforms such as ElevenLabs Agents handle transport, telephony and turn-taking), then add their own prompt, tools and logging.
Latency budget
Measure each segment: user stops speaking, turn detected, first model audio, tool round trips. Keep a latency budget per turn and test on mobile networks. Short first responses and parallel tool calls reduce perceived delay more than any single model change.
Compliance for voice agents
Disclose automation at the start of calls. Record and transcribe only with appropriate notice or consent under local law, and keep retention short. For outbound calling, check telemarketing and AI-voice rules in each market (for example rules in the US on AI-generated voices in robocalls, and consumer protection rules in the UK, UAE and Saudi Arabia). Offer a route to a human.
How to measure success
Track task completion, turn latency (median and slowest few percent), interruption handling, transfer rate, and caller satisfaction. Review a sample of recorded calls weekly with the prompt open beside you; most improvements come from small prompt and tool fixes found this way.
Key takeaways
- Voice agents use either an STT-LLM-TTS pipeline or a realtime speech-to-speech model; each has trade-offs.
- Turn detection, barge-in, filler phrases during tool calls and transport choices shape the experience.
- Voice prompts add rules for how things sound: short turns, spoken numbers, confirmation before actions.
- Disclose automation, respect recording and telemarketing rules, measure latency per segment and review calls.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Write a voice-agent system prompt for a real booking or support use case with style, task and boundary sections, then test five calls and log the latency of each segment.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.