Voice AI & Conversational AgentsPlatforms and stacks · Lesson 3 of 17

ElevenLabs voices: TTS models, voice design and cloning

Article · 15 min · 9 min lecture

Video lecture

ElevenLabs voices: TTS models, voice design and cloning

13 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 13

The voice is the brand

  • Model families and when to use them
  • Settings that matter
  • Four ways to get a voice
  • Writing for the ear

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The voice is the brand

In a voice agent, the voice is your front desk, your salesperson and your brand ambassador. It shapes trust in the first two seconds. This lesson focuses on ElevenLabs, one of the most widely used voice platforms, with notes on alternatives. Product names and model lineups change often; check the current models page before you build.

Model families (as of September 2026)

Model (API id)Best forNotes
Eleven v3 (eleven_v3)Expressive narration, dialogue, audiobooks, ads70+ languages; supports audio tags in square brackets such as [whispers], [laughs], [sighs] to direct delivery; multi-speaker dialogue
Eleven Multilingual v2 (eleven_multilingual_v2)Stable, high-quality long-form29 languages; consistent and reliable
Eleven Flash v2.5 (eleven_flash_v2_5)Real-time agentsUltra-low latency, 32 languages, lower cost per character
Eleven Turbo v2.5 (eleven_turbo_v2_5)Balance of quality and speed32 languages

Rule of thumb: Flash for live conversation, v3 or Multilingual v2 for produced content (course lectures, podcasts, ads). Conversational agents on ElevenAgents choose from the conversational TTS models the platform lists.

Speech-to-text: Scribe (scribe_v2, and a realtime variant) transcribes with timestamps, speaker diarization and audio-event tagging, useful for call analysis and captions.

Voice settings that matter

  • Stability: higher is more consistent and calm; lower is more expressive and variable. Agents usually want moderate-to-high stability.
  • Similarity boost: how closely to match the original voice; too high can reproduce recording artifacts.
  • Style (on some models): exaggerates the speaker's style; raises latency and variability.
  • Speed: slightly slower speech helps on phone lines and for non-native listeners.
  • Text normalization and pronunciation dictionaries: control how numbers, dates, acronyms and brand names are read. You can attach pronunciation dictionaries with alias or phoneme rules.

Four ways to get a voice

  1. Voice Library: choose from pre-made and community voices, with usage terms per voice.
  2. Voice design (text-to-voice): describe a voice ("warm, mid-30s female voice, Gulf Arabic accent, calm and professional, studio quality") and generate previews; save the one you like. No real person's voice is copied.
  3. Instant Voice Cloning (IVC): a quick clone from short samples; good for prototypes.
  4. Professional Voice Cloning (PVC): trained on longer, clean recordings for higher fidelity; platforms require verification that the voice belongs to you or that you have permission.

Cloning a real person's voice requires their informed, written consent (Module 5). Voice design is often the safer route for a brand voice.

Writing for the ear

TTS reads what you give it. Good scripts for voice:

  • Short sentences. One idea each.
  • Numbers written the way they should be said when ambiguous ("two thirty p m", "zero three zero zero").
  • Avoid symbols and abbreviations unless your normalization handles them.
  • Use punctuation for rhythm: commas for short pauses, full stops for longer ones. Some models support explicit break tags; v3 responds to audio tags and punctuation.

Worked example: a voice for a Karachi fintech's support agent

Requirements: Urdu and English, trustworthy, not robotic, fast. Process:

  1. Designed three voices with text-to-voice prompts; ran blind preference tests with 15 staff and 10 customers reading the same five lines, including an Urdu-English code-switched line ("Aap ka OTP 4 digits ka hai.").
  2. Chose one voice; tested it on Flash v2.5 for the live agent and Multilingual v2 for IVR prompts and marketing audio.
  3. Added a pronunciation dictionary for the brand name and product terms; set text normalization for amounts in PKR.
  4. Stability set moderately high; speed slightly below default for phone clarity.

Hands-on: generate, stream and save audio with the Python SDK

import os
from elevenlabs.client import ElevenLabs

client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
VOICE_ID = os.environ["BRAND_VOICE_ID"]  # from your Voice Library or a designed voice

# Produced content: expressive model with audio tags
audio = client.text_to_speech.convert(
    voice_id=VOICE_ID,
    model_id="eleven_v3",
    text="[warmly] Welcome to lesson three. [pause] Today, we build a voice people trust.",
    output_format="mp3_44100_128",
)
with open("intro.mp3", "wb") as f:
    for chunk in audio:
        f.write(chunk)

# Live-style streaming: low-latency model
stream = client.text_to_speech.stream(
    voice_id=VOICE_ID,
    model_id="eleven_flash_v2_5",
    text="Thanks for calling. How can I help today?",
)
with open("greeting.mp3", "wb") as f:
    for chunk in stream:
        if isinstance(chunk, bytes):
            f.write(chunk)

Handle errors (rate limits, invalid voice IDs) with retries and clear logging, and never commit API keys. Check the current docs for audio tag support, which varies by model.

Alternatives worth knowing

Other capable TTS providers include OpenAI (voices in its speech and realtime models), Google Cloud Text-to-Speech, Microsoft Azure AI Speech, Amazon Polly, Cartesia and Deepgram. Compare on your languages, latency from your region, voice quality with your script, licensing, and data terms.

Pitfalls

  • Choosing a voice from a 10-second demo instead of testing your real script, languages and phone-line audio.
  • Maximizing expressiveness for a support agent; callers want calm and clear.
  • Cloning a colleague's voice "just for a test" without written consent.

Key takeaways

  • Use low-latency models (e.g. Eleven Flash v2.5) for live agents and expressive models (Eleven v3, Multilingual v2) for produced content; check current model lists.
  • Tune stability, similarity and speed, and use text normalization and pronunciation dictionaries for numbers and brand terms.
  • Get voices from a library, voice design, or cloning; cloning a real person requires verified, written consent.
  • Write for the ear: short sentences, spelled-out ambiguous numbers, punctuation for rhythm; test real scripts on real phone audio.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which ElevenLabs model family is the usual choice for a live phone agent?
  2. Your agent mispronounces your brand name. What is the most direct fix?
  3. What is the safest route to a unique brand voice without copying a real person?

Put it into practice

Design two candidate voices from text descriptions, generate your agent's first three lines with each, play them through a phone speaker, and run a quick trust preference test with five people.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.