Voice AI & Conversational AgentsVoice agent foundations · Lesson 1 of 17

How voice agents work: cascaded pipelines vs speech-to-speech

Article · 15 min · 9 min lecture

Video lecture

How voice agents work: cascaded pipelines vs speech-to-speech

13 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 13

Voice AI & Conversational Agents

  • From IVR menus to real conversations
  • Design, build, test, launch
  • Architecture shapes everything

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What changed

For decades, "voice automation" meant IVR menus: press 1 for sales. Speech recognition could catch a few keywords, and everything else failed. Three things changed in 2023 to 2026: speech recognition became accurate across accents and noisy lines, large language models (LLMs) became good at open conversation and tool use, and speech synthesis became natural enough that many callers cannot immediately tell. Add streaming at every step, and you get voice agents: software that holds a spoken conversation, understands intent, looks things up, takes actions and hands off to humans.

The two architectures

1. Cascaded (pipeline) architecture: ASR -> LLM -> TTS

Caller audio -> [VAD / turn detection] -> [ASR: speech-to-text] -> text
     -> [LLM + tools + knowledge] -> response text
     -> [TTS: text-to-speech] -> agent audio -> caller
  • ASR (automatic speech recognition, also called STT) turns audio into text, ideally streaming partial transcripts as the caller speaks.
  • VAD (voice activity detection) and a turn-detection model decide when the caller has finished.
  • The LLM decides what to say and which tools to call.
  • TTS turns text into audio, streaming the first audio chunk as soon as possible.

Strengths: you can pick the best component for each job (a strong Urdu ASR, your preferred LLM, a brand voice), inspect text at every step, apply text-based guardrails, and swap parts independently. Weaknesses: latency adds up across hops, and emotion or tone in the caller's voice is mostly lost when converted to text.

2. Speech-to-speech (realtime, "native audio") architecture

A single multimodal model takes audio in and produces audio out directly, often with built-in turn detection. Examples include OpenAI's Realtime API with its gpt-realtime model family and Google's Gemini Live API with native-audio models. (Model names change frequently; check each provider's current model list.)

Strengths: lower latency potential, better handling of prosody (tone, hesitation, laughter), more natural interruptions. Weaknesses: less control over the exact voice (you choose from the provider's voices rather than any cloned brand voice in most cases), harder to insert text guardrails mid-stream, vendor lock-in, and audio-token costs.

Hybrid designs are common: a speech-to-speech model for conversation with text-based tool calls, or a cascaded stack using an LLM that accepts audio directly.

The platform layer

Most teams do not wire components by hand. They use a voice agent platform that bundles turn-taking, ASR, LLM orchestration, TTS, telephony, tools, knowledge bases, testing and analytics:

TypeExamples (verify current offerings)When to use
Integrated agent platformsElevenLabs' ElevenAgents (formerly "Conversational AI"), and orchestration platforms such as Vapi and RetellFastest path to production; configuration over code
Model-provider realtime APIsOpenAI Realtime API (WebRTC, WebSocket, SIP), Gemini Live APIYou want speech-to-speech and control in code
Open-source frameworksLiveKit Agents, PipecatFull control, self-hosting, custom pipelines
Component vendorsSTT from ElevenLabs (Scribe), Deepgram, AssemblyAI; TTS from ElevenLabs, Cartesia, Deepgram and othersBuilding your own cascade

Anatomy of a single turn (cascaded)

  1. Caller speaks: "Hi, can I move my appointment to Thursday afternoon?"
  2. Streaming ASR produces partial text while they speak.
  3. Turn detector decides the caller has finished (silence plus semantic cues).
  4. LLM receives the transcript plus conversation history and system prompt; it calls the check_availability tool with {"date": "Thursday", "period": "afternoon"}.
  5. While the tool runs, the agent says a short filler ("Let me check Thursday for you.").
  6. Tool returns slots; LLM writes: "I have 2:30 or 4 pm on Thursday. Which works better?"
  7. TTS streams audio; first audio plays well before the full sentence is synthesized.

Worked example: choosing an architecture

ScenarioChoiceWhy
Dental clinic in Dubai, Arabic and English, needs its own branded voice and strict booking rulesCascaded via an agent platformBrand voice, per-component language choice, text guardrails
Language-learning app practicing casual English conversationSpeech-to-speechNatural prosody and interruptions matter most
Bank in Riyadh with strict data residencyCascaded, possibly self-hosted componentsControl over where audio and transcripts are processed
Startup MVP to test demand in two weeksIntegrated platformSpeed to market

Hands-on: a minimal cascaded turn in Python

This non-streaming example shows the three hops clearly. Production agents stream every step, which platforms handle for you.

import os
from elevenlabs.client import ElevenLabs
from openai import OpenAI

eleven = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
llm = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

def one_turn(audio_path: str, voice_id: str) -> bytes:
    # 1) ASR: speech to text
    with open(audio_path, "rb") as f:
        stt = eleven.speech_to_text.convert(model_id="scribe_v2", file=f)
    user_text = stt.text

    # 2) LLM: decide what to say
    reply = llm.responses.create(
        model=os.environ.get("LLM_MODEL", "gpt-5-mini"),
        instructions="You are a concise, friendly receptionist. Reply in one or two short spoken sentences.",
        input=user_text,
    ).output_text

    # 3) TTS: text to speech (low-latency model for conversation)
    audio_chunks = eleven.text_to_speech.convert(
        voice_id=voice_id, text=reply, model_id="eleven_flash_v2_5", output_format="mp3_44100_128")
    return b"".join(audio_chunks)

Time each hop with time.perf_counter() and you will see why streaming and component choice matter; the next lesson covers latency budgets.

Pitfalls

  • Choosing an architecture by demo impressiveness instead of requirements (voice, languages, control, compliance).
  • Forgetting that a voice agent is also an AI system with legal duties (disclosure, recording consent, data protection), covered in Module 5.
  • Treating the LLM prompt as the whole product; turn-taking and latency often matter more to callers.

Key takeaways

  • Voice agents combine turn detection, ASR, an LLM with tools and knowledge, and TTS, streaming at each step.
  • Cascaded pipelines give control, brand voices and inspectable text; speech-to-speech models give lower-latency, more natural prosody but less control.
  • Integrated platforms (e.g. ElevenAgents), realtime model APIs and open-source frameworks (LiveKit Agents, Pipecat) are the main build paths.
  • Choose architecture by requirements: voice, languages, control, compliance and time to market.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A Dubai clinic needs its own branded Arabic and English voice and strict, auditable booking rules. Which architecture fits best?
  2. What is a main advantage of speech-to-speech models?
  3. During a tool call that takes a second or two, what should a well-designed agent usually do?

Put it into practice

Run the one-turn Python example with a short recorded question. Time each hop (ASR, LLM, TTS) and note which is slowest.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.