Voice AI & Conversational AgentsVoice agent foundations · Lesson 1 of 17
How voice agents work: cascaded pipelines vs speech-to-speech
Video lecture
How voice agents work: cascaded pipelines vs speech-to-speech
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Voice AI & Conversational Agents
Call a business today and there's a real chance a voice agent picks up. Not a press one for sales menu, but something that understands you, checks a calendar, and books you in. In this course you'll learn to design, build, test and launch voice agents that people actually like talking to, legally and ethically. We'll start with how they work, because the architecture you pick shapes everything: cost, latency, voice quality, languages and control.
0:33 Why architecture first
Why start with architecture? Because it decides almost everything that follows: how fast the agent can respond, which voices you can use, how many languages you can support well, how easily you can add guardrails, what it costs per minute, and how locked in you are to one vendor. Teams that pick an architecture by accident often rebuild six months later. A few minutes of deliberate choice now saves that rebuild.
1:04 What changed
What changed? Three things, roughly between twenty twenty-three and twenty twenty-six. Speech recognition got accurate across accents and noisy phone lines. Large language models got good at open conversation and at calling tools. And speech synthesis got natural enough that many callers can't immediately tell. Add streaming at every step, and you get voice agents: software that holds a spoken conversation, understands intent, looks things up, takes actions, and hands off to a human when needed.
1:37 Analogy: relay team vs interpreter
Here's an analogy for the two architectures. A cascaded agent is like a relay team. One runner listens and writes down what was said, the next thinks and writes a reply, and the last reads it aloud. Each runner can be the best in the world at their leg, and you can swap any of them, but every handover costs time. A speech to speech model is like one talented interpreter who listens and responds in a single flow. Smoother and faster, but you get their voice and their style, and it's harder to check their notes mid-sentence.
2:20 Cascaded pipeline
The first architecture is the cascade, or pipeline. Voice activity detection and a turn detector decide when the caller has finished. Automatic speech recognition turns their words into text. A language model decides what to say and which tools to call. And text to speech turns the reply back into audio. The big advantage is control: you can pick the best component for each job, like a strong Urdu recognizer and your own brand voice, inspect the text at every step, and swap parts. The costs: latency adds up across hops, and the caller's tone mostly disappears when audio becomes text.
3:04 Speech-to-speech
The second architecture is speech to speech. One multimodal model takes audio in and produces audio out, usually with turn detection built in. OpenAI's Realtime API and Google's Gemini Live API are examples. The upside: potentially lower latency, and much better handling of tone, hesitation and interruptions. The downsides: you usually choose from the provider's voices rather than your own cloned brand voice, it's harder to insert text guardrails mid-stream, and you're more tied to one vendor. Many teams end up with hybrids, mixing both approaches.
3:41 The platform layer
In practice, most teams don't wire components by hand. They use a platform. Integrated agent platforms, like ElevenLabs' ElevenAgents, or orchestration platforms such as Vapi and Retell, bundle turn-taking, speech, the language model, telephony, tools, knowledge bases, testing and analytics. Model providers offer realtime APIs for speech to speech in code. Open-source frameworks like LiveKit Agents and Pipecat give full control and self-hosting. And component vendors sell speech recognition and synthesis separately if you want to build your own cascade.
4:16 One turn, step by step
Let's follow one turn. The caller says: can I move my appointment to Thursday afternoon? Streaming recognition produces partial text while they speak. The turn detector decides they're done. The language model reads the transcript, the history and its instructions, and calls a check availability tool for Thursday afternoon. While the tool runs, the agent says: let me check Thursday for you. The tool returns slots, and the model replies: I have two thirty or four p m on Thursday, which works better? Speech synthesis streams it, so audio starts playing before the whole sentence is even generated.
4:59 Choosing an architecture
So how do you choose? A dental clinic in Dubai that needs Arabic and English, its own branded voice and strict booking rules usually fits a cascaded platform. A language-learning app practicing casual conversation benefits from speech to speech, because natural rhythm matters most. A bank in Riyadh with strict data residency may want a cascade with components it can host or control. And a startup testing demand in two weeks should use an integrated platform. Choose by requirements: voice, languages, control, compliance and speed, not by which demo sounded coolest.
5:39 Example 2: Lahore yoga studio
A simple example to make it concrete. A yoga studio in Lahore wants a phone line that answers three questions: class times, prices and location. No bookings, no payments. They don't need a custom build at all. An integrated platform agent with a short prompt, a small knowledge base with the timetable and prices, and a friendly designed voice does the job in an afternoon. The architecture choice barely matters here. It starts to matter when you need specific voices, strict rules, many languages or deep integrations.
6:17 Common mistakes
Three common mistakes when starting out. Choosing an architecture because a demo sounded impressive, instead of checking your requirements for voice, languages, control and compliance. Forgetting that a voice agent is also an AI system with legal duties, like disclosure, recording notices and data protection. And treating the prompt as the whole product. Callers judge you on timing, turn-taking and how gracefully the agent recovers, often more than on the words themselves.
6:48 Watch me do it: start the agent config
Watch me do it. I'm starting the agent config file we'll grow through this whole course. I create a file called nova agent dot json for Nova Dental. First, I write the architecture decision as a comment block, because future me will want to know why. Requirements: English and Arabic callers, a calm branded voice, strict booking rules, phone and web. Decision: cascaded pipeline on an agent platform. Reason: brand voice and text-level guardrails matter more than the last bit of natural prosody. Next, I add the three layers as empty sections: speech recognition, with a note to test Arabic dialects; language model, with a note saying fast tier for routine turns; and text to speech, with a note saying low-latency model, voice to be designed in lesson three. Then I add a section called turn taking, empty for now. Then I run the one-turn Python script from the lesson text on a recording of me asking, can I move my appointment to Thursday afternoon. The terminal prints three timings: speech recognition, the model and synthesis. I paste those numbers into the config comments as our first baseline. The slowest hop here is the model, so I note that too. That's our starting point: a written decision, an empty skeleton, and real numbers. Every lesson from now on fills in one more part of this file.
8:26 Recap and next step
Recap. Voice agents combine turn detection, speech recognition, a language model with tools, and speech synthesis, either as a cascade or as one speech to speech model. Cascades give control and brand voice. Speech to speech gives natural flow. Platforms bundle it all. Your next step: run the short Python example in the lesson text, one turn through recognition, a language model and synthesis, and time each hop. Those numbers set up the next lesson on latency.
What changed
For decades, "voice automation" meant IVR menus: press 1 for sales. Speech recognition could catch a few keywords, and everything else failed. Three things changed in 2023 to 2026: speech recognition became accurate across accents and noisy lines, large language models (LLMs) became good at open conversation and tool use, and speech synthesis became natural enough that many callers cannot immediately tell. Add streaming at every step, and you get voice agents: software that holds a spoken conversation, understands intent, looks things up, takes actions and hands off to humans.
The two architectures
1. Cascaded (pipeline) architecture: ASR -> LLM -> TTS
Caller audio -> [VAD / turn detection] -> [ASR: speech-to-text] -> text
-> [LLM + tools + knowledge] -> response text
-> [TTS: text-to-speech] -> agent audio -> caller- ASR (automatic speech recognition, also called STT) turns audio into text, ideally streaming partial transcripts as the caller speaks.
- VAD (voice activity detection) and a turn-detection model decide when the caller has finished.
- The LLM decides what to say and which tools to call.
- TTS turns text into audio, streaming the first audio chunk as soon as possible.
Strengths: you can pick the best component for each job (a strong Urdu ASR, your preferred LLM, a brand voice), inspect text at every step, apply text-based guardrails, and swap parts independently. Weaknesses: latency adds up across hops, and emotion or tone in the caller's voice is mostly lost when converted to text.
2. Speech-to-speech (realtime, "native audio") architecture
A single multimodal model takes audio in and produces audio out directly, often with built-in turn detection. Examples include OpenAI's Realtime API with its gpt-realtime model family and Google's Gemini Live API with native-audio models. (Model names change frequently; check each provider's current model list.)
Strengths: lower latency potential, better handling of prosody (tone, hesitation, laughter), more natural interruptions. Weaknesses: less control over the exact voice (you choose from the provider's voices rather than any cloned brand voice in most cases), harder to insert text guardrails mid-stream, vendor lock-in, and audio-token costs.
Hybrid designs are common: a speech-to-speech model for conversation with text-based tool calls, or a cascaded stack using an LLM that accepts audio directly.
The platform layer
Most teams do not wire components by hand. They use a voice agent platform that bundles turn-taking, ASR, LLM orchestration, TTS, telephony, tools, knowledge bases, testing and analytics:
| Type | Examples (verify current offerings) | When to use |
|---|---|---|
| Integrated agent platforms | ElevenLabs' ElevenAgents (formerly "Conversational AI"), and orchestration platforms such as Vapi and Retell | Fastest path to production; configuration over code |
| Model-provider realtime APIs | OpenAI Realtime API (WebRTC, WebSocket, SIP), Gemini Live API | You want speech-to-speech and control in code |
| Open-source frameworks | LiveKit Agents, Pipecat | Full control, self-hosting, custom pipelines |
| Component vendors | STT from ElevenLabs (Scribe), Deepgram, AssemblyAI; TTS from ElevenLabs, Cartesia, Deepgram and others | Building your own cascade |
Anatomy of a single turn (cascaded)
- Caller speaks: "Hi, can I move my appointment to Thursday afternoon?"
- Streaming ASR produces partial text while they speak.
- Turn detector decides the caller has finished (silence plus semantic cues).
- LLM receives the transcript plus conversation history and system prompt; it calls the
check_availabilitytool with{"date": "Thursday", "period": "afternoon"}. - While the tool runs, the agent says a short filler ("Let me check Thursday for you.").
- Tool returns slots; LLM writes: "I have 2:30 or 4 pm on Thursday. Which works better?"
- TTS streams audio; first audio plays well before the full sentence is synthesized.
Worked example: choosing an architecture
| Scenario | Choice | Why |
|---|---|---|
| Dental clinic in Dubai, Arabic and English, needs its own branded voice and strict booking rules | Cascaded via an agent platform | Brand voice, per-component language choice, text guardrails |
| Language-learning app practicing casual English conversation | Speech-to-speech | Natural prosody and interruptions matter most |
| Bank in Riyadh with strict data residency | Cascaded, possibly self-hosted components | Control over where audio and transcripts are processed |
| Startup MVP to test demand in two weeks | Integrated platform | Speed to market |
Hands-on: a minimal cascaded turn in Python
This non-streaming example shows the three hops clearly. Production agents stream every step, which platforms handle for you.
import os
from elevenlabs.client import ElevenLabs
from openai import OpenAI
eleven = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
llm = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
def one_turn(audio_path: str, voice_id: str) -> bytes:
# 1) ASR: speech to text
with open(audio_path, "rb") as f:
stt = eleven.speech_to_text.convert(model_id="scribe_v2", file=f)
user_text = stt.text
# 2) LLM: decide what to say
reply = llm.responses.create(
model=os.environ.get("LLM_MODEL", "gpt-5-mini"),
instructions="You are a concise, friendly receptionist. Reply in one or two short spoken sentences.",
input=user_text,
).output_text
# 3) TTS: text to speech (low-latency model for conversation)
audio_chunks = eleven.text_to_speech.convert(
voice_id=voice_id, text=reply, model_id="eleven_flash_v2_5", output_format="mp3_44100_128")
return b"".join(audio_chunks)Time each hop with time.perf_counter() and you will see why streaming and component choice matter; the next lesson covers latency budgets.
Pitfalls
- Choosing an architecture by demo impressiveness instead of requirements (voice, languages, control, compliance).
- Forgetting that a voice agent is also an AI system with legal duties (disclosure, recording consent, data protection), covered in Module 5.
- Treating the LLM prompt as the whole product; turn-taking and latency often matter more to callers.
Key takeaways
- Voice agents combine turn detection, ASR, an LLM with tools and knowledge, and TTS, streaming at each step.
- Cascaded pipelines give control, brand voices and inspectable text; speech-to-speech models give lower-latency, more natural prosody but less control.
- Integrated platforms (e.g. ElevenAgents), realtime model APIs and open-source frameworks (LiveKit Agents, Pipecat) are the main build paths.
- Choose architecture by requirements: voice, languages, control, compliance and time to market.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Run the one-turn Python example with a short recorded question. Time each hop (ASR, LLM, TTS) and note which is slowest.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.