Voice AI & Conversational AgentsPlatforms and stacks · Lesson 3 of 17
ElevenLabs voices: TTS models, voice design and cloning
Video lecture
ElevenLabs voices: TTS models, voice design and cloning
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 The voice is the brand
Close your eyes and imagine calling your bank and hearing a voice that sounds tinny, rushed and robotic. Now imagine a calm, warm, clear voice. Same words, same information, completely different trust. In a voice agent, the voice is your front desk and your brand. In this lesson, you'll learn the ElevenLabs model families and when to use each, the voice settings that matter, four ways to get a voice, how to write for the ear, and how to generate and stream audio in code.
0:37 Why the voice matters
Why spend time on the voice? Because callers form an impression in the first two seconds, before the agent has said anything useful. A voice that sounds rushed, robotic or mismatched to your brand lowers trust, and a trusted voice makes people more patient with small mistakes. The voice also affects comprehension on phone lines and for non-native listeners. It's one of the cheapest, highest-impact decisions you'll make.
1:07 ElevenLabs TTS families
As of September twenty twenty-six, ElevenLabs offers four main text to speech families. Eleven v3 is the expressive one, for narration, dialogue and ads, with more than seventy languages and audio tags, little stage directions in square brackets, like whispers or laughs. Multilingual v2 is stable and reliable for long-form content in twenty-nine languages. Flash v2.5 is built for real-time agents, with ultra-low latency. And Turbo v2.5 balances quality and speed. The rule of thumb: Flash for live conversation, v3 or Multilingual v2 for produced content. Always check the current models page, because lineups change.
1:48 Analogy: casting performers
An analogy: choosing a voice model is like choosing a performer for different jobs. For a live radio call-in show, you want someone quick, clear and unflappable, that's Flash. For an audiobook, you want someone expressive who can whisper, laugh and build suspense, that's Eleven v3 with audio tags. For a long documentary series, you want someone rock-steady for hours, that's Multilingual v2. Same talent agency, different performers, each cast for the role.
2:20 Settings that matter
A few voice settings make a big difference. Stability: higher is calmer and more consistent, lower is more expressive but variable. Agents usually want it moderate to high. Similarity boost controls how closely the output matches the original voice, but too high can copy recording flaws. Speed: a touch slower helps on phone lines and for non-native listeners. And text normalization plus pronunciation dictionaries control how numbers, dates, acronyms and brand names are read. That's how you stop your agent mangling your own company name.
2:57 Four ways to get a voice
There are four ways to get a voice. Pick one from the Voice Library, checking each voice's usage terms. Design one from a text description, for example a warm, mid-thirties voice with a Gulf Arabic accent, calm and professional, then choose from previews. No real person is copied. Make an instant clone from short samples, which is fine for prototypes. Or train a professional clone on longer, clean recordings, where platforms require verification that you own the voice or have permission. Cloning any real person needs their informed, written consent. For a brand voice, voice design is often the safest route.
3:41 Writing for the ear
Text to speech reads exactly what you give it, so write for the ear. Keep sentences short, one idea each. Write numbers the way they should be spoken when they could be ambiguous, like two thirty p m or zero three zero zero. Avoid symbols and abbreviations unless normalization handles them. And use punctuation for rhythm: commas for short pauses, full stops for longer ones. With Eleven v3, audio tags and punctuation shape the delivery. Test your real script, not a demo sentence.
4:17 Worked example: Karachi fintech voice
Here's how a Karachi fintech chose the voice for its support agent. They needed Urdu and English, trustworthy, not robotic, and fast. They designed three voices from text prompts and ran blind preference tests with staff and customers using five real lines, including a code-switched one mixing Urdu and English. They picked a winner, used Flash for the live agent and Multilingual v2 for recorded prompts and marketing audio, added a pronunciation dictionary for the brand and product names, set normalization for rupee amounts, and nudged speed slightly below default for phone clarity.
4:57 In code
In code, the Python SDK makes this simple. For produced content, call text to speech convert with the v3 model, your voice I D and your script, including audio tags, and write the chunks to a file. For live-style output, call the stream method with the Flash model and write or play chunks as they arrive. Keep your API key in an environment variable, handle rate limits with retries, and log errors clearly. The full example is in the lesson text. And do compare alternatives, like OpenAI, Google, Microsoft, Amazon, Cartesia and Deepgram, on your languages, latency, quality, licensing and data terms.
5:41 Example 2: Manchester podcaster
A simple example. A solo podcaster in Manchester wants an intro and outro in a consistent voice every week, without recording them herself. She designs a warm, confident voice from a text description, writes a thirty-second intro script, generates it with Multilingual v2 for stability, and saves the voice settings in a small config file. Each week only the episode title changes, so she regenerates one line. It sounds identical every time, and she lists the intro as AI-voiced in her show notes.
6:17 Common mistakes
Three mistakes with voices. Choosing from a ten-second demo instead of testing your real script, in your languages, over a real phone line. Maximizing expressiveness for a support agent, when callers want calm and clear, not dramatic. And cloning a colleague's or celebrity's voice for a quick test without written consent. The test becomes a file, the file gets shared, and suddenly you have a real problem.
6:46 Watch me do it: design and configure the voice
Watch me do it. I open the voice design tool and type a description: a warm, calm female voice in her thirties, clear and professional, suitable for a dental clinic receptionist, neutral international English with the ability to speak Arabic naturally. I generate previews using our real first line: hi, I'm Aria, Nova Dental's AI assistant. How can I help today? I listen to three previews on my phone's speaker, not my headphones. Preview two sounds warmest without sounding sleepy, so I save it and copy the voice I D into the agent config's voice section. Next, model: for the live agent I set the low-latency Flash model. I set stability moderately high and speed slightly below default. Then I create a pronunciation dictionary with three rules: Nova Dental, the Marina branch name, and the word hygienist, which the voice slightly rushed. I attach the dictionary to the config. Finally, I generate the same greeting with Eleven v3 for our on-hold message, adding an audio tag for a warm tone, and save it as an MP3. I play both versions to two colleagues and ask one question: would you trust this receptionist? Both say yes to the live version, and one asks to slow the on-hold message a touch, so I do.
8:18 Recap and next step
Recap. Pick the model for the job: Flash for live conversation, v3 or Multilingual v2 for produced audio. Tune stability, similarity and speed, and control pronunciation. Get a voice from the library, by design, or by cloning with verified consent. Write for the ear. Your next step: design two candidate voices from text descriptions, generate your agent's first three lines with each, play them over a phone line or phone speaker, and ask five people which they trust more.
The voice is the brand
In a voice agent, the voice is your front desk, your salesperson and your brand ambassador. It shapes trust in the first two seconds. This lesson focuses on ElevenLabs, one of the most widely used voice platforms, with notes on alternatives. Product names and model lineups change often; check the current models page before you build.
Model families (as of September 2026)
| Model (API id) | Best for | Notes |
|---|---|---|
Eleven v3 (eleven_v3) | Expressive narration, dialogue, audiobooks, ads | 70+ languages; supports audio tags in square brackets such as [whispers], [laughs], [sighs] to direct delivery; multi-speaker dialogue |
Eleven Multilingual v2 (eleven_multilingual_v2) | Stable, high-quality long-form | 29 languages; consistent and reliable |
Eleven Flash v2.5 (eleven_flash_v2_5) | Real-time agents | Ultra-low latency, 32 languages, lower cost per character |
Eleven Turbo v2.5 (eleven_turbo_v2_5) | Balance of quality and speed | 32 languages |
Rule of thumb: Flash for live conversation, v3 or Multilingual v2 for produced content (course lectures, podcasts, ads). Conversational agents on ElevenAgents choose from the conversational TTS models the platform lists.
Speech-to-text: Scribe (scribe_v2, and a realtime variant) transcribes with timestamps, speaker diarization and audio-event tagging, useful for call analysis and captions.
Voice settings that matter
- Stability: higher is more consistent and calm; lower is more expressive and variable. Agents usually want moderate-to-high stability.
- Similarity boost: how closely to match the original voice; too high can reproduce recording artifacts.
- Style (on some models): exaggerates the speaker's style; raises latency and variability.
- Speed: slightly slower speech helps on phone lines and for non-native listeners.
- Text normalization and pronunciation dictionaries: control how numbers, dates, acronyms and brand names are read. You can attach pronunciation dictionaries with alias or phoneme rules.
Four ways to get a voice
- Voice Library: choose from pre-made and community voices, with usage terms per voice.
- Voice design (text-to-voice): describe a voice ("warm, mid-30s female voice, Gulf Arabic accent, calm and professional, studio quality") and generate previews; save the one you like. No real person's voice is copied.
- Instant Voice Cloning (IVC): a quick clone from short samples; good for prototypes.
- Professional Voice Cloning (PVC): trained on longer, clean recordings for higher fidelity; platforms require verification that the voice belongs to you or that you have permission.
Cloning a real person's voice requires their informed, written consent (Module 5). Voice design is often the safer route for a brand voice.
Writing for the ear
TTS reads what you give it. Good scripts for voice:
- Short sentences. One idea each.
- Numbers written the way they should be said when ambiguous ("two thirty p m", "zero three zero zero").
- Avoid symbols and abbreviations unless your normalization handles them.
- Use punctuation for rhythm: commas for short pauses, full stops for longer ones. Some models support explicit break tags; v3 responds to audio tags and punctuation.
Worked example: a voice for a Karachi fintech's support agent
Requirements: Urdu and English, trustworthy, not robotic, fast. Process:
- Designed three voices with text-to-voice prompts; ran blind preference tests with 15 staff and 10 customers reading the same five lines, including an Urdu-English code-switched line ("Aap ka OTP 4 digits ka hai.").
- Chose one voice; tested it on Flash v2.5 for the live agent and Multilingual v2 for IVR prompts and marketing audio.
- Added a pronunciation dictionary for the brand name and product terms; set text normalization for amounts in PKR.
- Stability set moderately high; speed slightly below default for phone clarity.
Hands-on: generate, stream and save audio with the Python SDK
import os
from elevenlabs.client import ElevenLabs
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
VOICE_ID = os.environ["BRAND_VOICE_ID"] # from your Voice Library or a designed voice
# Produced content: expressive model with audio tags
audio = client.text_to_speech.convert(
voice_id=VOICE_ID,
model_id="eleven_v3",
text="[warmly] Welcome to lesson three. [pause] Today, we build a voice people trust.",
output_format="mp3_44100_128",
)
with open("intro.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)
# Live-style streaming: low-latency model
stream = client.text_to_speech.stream(
voice_id=VOICE_ID,
model_id="eleven_flash_v2_5",
text="Thanks for calling. How can I help today?",
)
with open("greeting.mp3", "wb") as f:
for chunk in stream:
if isinstance(chunk, bytes):
f.write(chunk)Handle errors (rate limits, invalid voice IDs) with retries and clear logging, and never commit API keys. Check the current docs for audio tag support, which varies by model.
Alternatives worth knowing
Other capable TTS providers include OpenAI (voices in its speech and realtime models), Google Cloud Text-to-Speech, Microsoft Azure AI Speech, Amazon Polly, Cartesia and Deepgram. Compare on your languages, latency from your region, voice quality with your script, licensing, and data terms.
Pitfalls
- Choosing a voice from a 10-second demo instead of testing your real script, languages and phone-line audio.
- Maximizing expressiveness for a support agent; callers want calm and clear.
- Cloning a colleague's voice "just for a test" without written consent.
Key takeaways
- Use low-latency models (e.g. Eleven Flash v2.5) for live agents and expressive models (Eleven v3, Multilingual v2) for produced content; check current model lists.
- Tune stability, similarity and speed, and use text normalization and pronunciation dictionaries for numbers and brand terms.
- Get voices from a library, voice design, or cloning; cloning a real person requires verified, written consent.
- Write for the ear: short sentences, spelled-out ambiguous numbers, punctuation for rhythm; test real scripts on real phone audio.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Design two candidate voices from text descriptions, generate your agent's first three lines with each, play them through a phone speaker, and run a quick trust preference test with five people.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.