Multimodal & Reasoning Models in PracticeAudio: speech-to-text and text-to-speech · Lesson 7 of 17

Realtime voice agents: speech-to-speech in practice

Article · 13 min · 8 min lecture

Video lecture

Realtime voice agents: speech-to-speech in practice

11 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 11

Realtime voice agents

  • Two architectures
  • Building blocks
  • Prompting for voice
  • Sessions, latency and compliance

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

From pipelines to speech-to-speech

Voice agents used to be built as a chain: speech-to-text, then a language model, then text-to-speech. Each stage added delay, and tone of voice was lost in the middle. Realtime speech-to-speech models now take audio in and produce audio out directly, over a persistent connection (typically WebSocket or WebRTC). Examples include OpenAI's Realtime API with its realtime voice models and Google's Gemini Live API. The component pipeline is still widely used, especially when you need a particular STT or TTS vendor, maximum control or specific languages.

ApproachStrengthsTrade-offs
Component pipeline (STT, LLM, TTS)Choose best-in-class parts, easy to log text at each stage, flexible languages and voicesMore latency to manage, tone and emotion lost between stages
Speech-to-speech realtime modelLower latency, natural turn-taking, preserves tone, handles interruptionsFewer voice choices, less visibility into intermediate text, provider lock-in

The building blocks of any voice agent

  • Turn detection (voice activity detection): deciding when the user has finished speaking. Too eager and the agent interrupts; too slow and it feels laggy. Most realtime APIs offer server-side turn detection with tunable sensitivity.
  • Barge-in: letting the user interrupt the agent, which must stop speaking immediately and listen.
  • Tool calls mid-conversation: checking a booking system, then speaking the answer. Say a short filler ("Let me check that") while tools run.
  • Telephony and transport: browsers via WebRTC, phone lines via SIP or a telephony provider.

Prompting a voice agent

Voice prompts need everything a text system prompt needs, plus rules for how things sound:

You are the booking assistant for Harbour Dental Clinic in Dubai.
You speak on the phone with patients in English or Arabic; reply in the
language the caller uses.

Style:
- Warm, calm and concise. One or two short sentences per turn.
- Never read out lists longer than three items; offer to send details by SMS.
- Say times naturally ("ten thirty in the morning"), not "10:30".
- Confirm names by spelling back only if the caller spells them first.

Task:
- Help callers book, move or cancel appointments using the tools.
- Before booking, confirm the date, time and dentist in one sentence and
  wait for a clear "yes".

Boundaries:
- You are an automated assistant; say so at the start of every call.
- Do not give medical advice. For pain, swelling or emergencies, give the
  emergency number and offer to transfer to reception.
- If the caller is frustrated or asks for a person twice, transfer.

Hands-on: a realtime session configuration

Realtime APIs are configured at session start with instructions, voice, turn detection and tools. The shape below follows the OpenAI Realtime API's session update event; field names evolve, so confirm against the current reference before use.

import json, os
# Sent as the first event after opening the realtime WebSocket or WebRTC data channel
session_update = {
    "type": "session.update",
    "session": {
        "type": "realtime",
        "model": os.environ["REALTIME_MODEL"],
        "instructions": open("prompts/harbour_dental_voice.md", encoding="utf-8").read(),
        "audio": {
            "input": {"turn_detection": {"type": "server_vad"}},
            "output": {"voice": os.environ.get("REALTIME_VOICE", "marin")},
        },
        "tools": [{
            "type": "function", "name": "find_slots",
            "description": "Find free appointment slots for a dentist on a date (YYYY-MM-DD).",
            "parameters": {"type": "object", "properties": {
                "dentist": {"type": "string"}, "date": {"type": "string"}},
                "required": ["dentist", "date"]},
        }],
    },
}
payload = json.dumps(session_update)  # send over the open connection

Most teams start from the provider's official realtime examples or a voice-agent framework (open-source frameworks and hosted voice-agent platforms such as ElevenLabs Agents handle transport, telephony and turn-taking), then add their own prompt, tools and logging.

Latency budget

Measure each segment: user stops speaking, turn detected, first model audio, tool round trips. Keep a latency budget per turn and test on mobile networks. Short first responses and parallel tool calls reduce perceived delay more than any single model change.

Compliance for voice agents

Disclose automation at the start of calls. Record and transcribe only with appropriate notice or consent under local law, and keep retention short. For outbound calling, check telemarketing and AI-voice rules in each market (for example rules in the US on AI-generated voices in robocalls, and consumer protection rules in the UK, UAE and Saudi Arabia). Offer a route to a human.

How to measure success

Track task completion, turn latency (median and slowest few percent), interruption handling, transfer rate, and caller satisfaction. Review a sample of recorded calls weekly with the prompt open beside you; most improvements come from small prompt and tool fixes found this way.

Key takeaways

  • Voice agents use either an STT-LLM-TTS pipeline or a realtime speech-to-speech model; each has trade-offs.
  • Turn detection, barge-in, filler phrases during tool calls and transport choices shape the experience.
  • Voice prompts add rules for how things sound: short turns, spoken numbers, confirmation before actions.
  • Disclose automation, respect recording and telemarketing rules, measure latency per segment and review calls.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which is a key advantage of a speech-to-speech realtime model over an STT-LLM-TTS pipeline?
  2. Your voice agent keeps interrupting callers mid-sentence. What should you adjust first?
  3. Which instruction suits a voice agent's system prompt but would be unusual in a text chatbot?

Put it into practice

Write a voice-agent system prompt for a real booking or support use case with style, task and boundary sections, then test five calls and log the latency of each segment.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.