AI Video & Voice Production: ElevenLabs, Veo, Runway and MoreAI voice: text-to-speech, cloning and dubbing · Lesson 1 of 16

How AI voice works: text-to-speech in plain language

Video lesson · 9 min · 8 min lecture

Video lecture

How AI voice works: text-to-speech in plain language

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

How AI voice works

  • Neural text-to-speech in plain English
  • Reading a model menu
  • Testing voices properly
  • Writing a voice spec

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

From robotic to realistic

Text-to-speech (TTS) converts written text into spoken audio. Earlier systems stitched together recorded sound fragments, which is why they sounded choppy. Modern neural TTS models are trained on large amounts of recorded speech and generate audio directly, learning pronunciation, intonation, pacing, breath and emotional color. The best current voices are hard to distinguish from human narration in short clips.

Leading voice platforms such as ElevenLabs typically offer a family of related capabilities:

  • Text-to-speech: type or paste a script, pick a voice, generate audio.
  • Speech-to-speech (voice changing): record yourself with the timing and emotion you want, then convert it into another voice while keeping your performance.
  • Voice design: create a new synthetic voice from a text description.
  • Voice cloning: create a voice that replicates a specific real person, with that person's consent.
  • Dubbing and translation: translate audio or video into other languages while preserving voice characteristics.
  • Transcription (speech-to-text), sound effects and conversational voice agents as related tools.

Features, model names and plans change frequently, so check the current documentation of whichever tool you use. The section "The 2026 model line-up" below shows how to read a model menu.

Key concepts

Voice types

  • Library or stock voices: pre-made voices you can use under the platform's license. Fast, low risk, but other brands may use the same voice.
  • Designed voices: generated from a description ("calm, deep, middle-aged male narrator with a neutral Gulf Arabic accent"). Unique to you, no real person's likeness involved.
  • Cloned voices: trained on recordings of a real person. Best for creators who want to scale their own voice, and only with explicit consent.

Models

Platforms usually offer several models that trade off quality, emotional range, latency and language coverage. A fast, low-latency model suits live agents; a more expressive model suits ads and storytelling. Multilingual models can speak many languages, but quality varies by language and accent. Always test in your target language, for example Urdu, Arabic or Hindi, before committing.

Settings you will meet

Exact names vary, but common controls include:

  • Stability: higher values give a more consistent, even delivery; lower values allow more variation and emotion, but also more unpredictability.
  • Similarity or clarity: how closely output sticks to the original voice's characteristics.
  • Style or expressiveness: exaggerates the voice's style; too much can sound theatrical.
  • Speed: delivery pace.

Start with default settings, change one at a time and compare.

The script is your direction

An AI narrator cannot ask what you meant. Your text is the direction:

  • Punctuation shapes pacing. Commas create short pauses; full stops create longer ones; dashes and ellipses add hesitation.
  • Sentence length shapes energy. Short sentences feel punchy. Long ones feel calm and explanatory.
  • Spell out tricky items. Write "twenty twenty-six" or "2026" consistently, and spell out brand names, acronyms and local names phonetically if mispronounced. Many tools support pronunciation dictionaries or phonetic hints for recurring terms.
  • Some models accept delivery cues in text, such as emotional or tonal tags, while others read everything literally. Test before relying on them.

Worked example

Nadia, a Karachi-based e-commerce marketer, needs English and Urdu voiceovers for 20 product videos a month. She tests three library voices and one designed voice across five sample scripts in both languages, rating naturalness, pronunciation of product names and consistency. The designed voice wins in English; a different library voice handles Urdu better. She creates a pronunciation list for the brand name and product terms, and documents the chosen voice, model and settings in a one-page "voice spec" so every video sounds consistent.

The 2026 model line-up (how to read it)

As of September 2026, ElevenLabs documents a small family of text-to-speech models, and the pattern generalizes to other vendors:

Model (API id)Best forTrade-off
Eleven v3 (eleven_v3)Expressive performances, ads, storytelling, multi-speaker dialogue; 70+ languagesMost emotional range; less predictable, so generate a few takes
Eleven Multilingual v2 (eleven_multilingual_v2)Long-form narration, courses, audiobooks; stable deliveryLess dramatic range than v3
Eleven Flash v2.5 (eleven_flash_v2_5)Real-time agents and previews; very low latencySlightly lower quality than the flagship models

Eleven v3 also understands inline audio tags in square brackets, such as [whispers], [excited] or [sighs], which act as stage directions. Older models read them literally, which is why you always test a cue before you rely on it. Model names change: before a project, open the vendor's models page and write the model id into your voice spec.

Hands-on: generate one line two ways

This short Python script uses the official elevenlabs SDK to render the same line with two models so you can compare them. Store your key in an environment variable, never in code.

# pip install elevenlabs python-dotenv
import os
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs

load_dotenv()                      # reads ELEVENLABS_API_KEY from .env
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

VOICE_ID = os.environ.get("VOICE_ID", "JBFqnCBsd6RMkjVDRZzb")  # any voice from your library
LINE = "Your first ten orders are on us. [pause] Seriously, that's the offer."

for model in ["eleven_multilingual_v2", "eleven_v3"]:
    try:
        audio = client.text_to_speech.convert(
            text=LINE, voice_id=VOICE_ID, model_id=model,
            output_format="mp3_44100_128",
        )
        with open(f"take_{model}.mp3", "wb") as f:
            for chunk in audio:          # the SDK returns the audio as a stream of bytes
                f.write(chunk)
        print("saved", model)
    except Exception as err:             # quota, invalid voice id, network
        print("failed", model, err)

Listen to both takes. With Multilingual v2 you will probably hear the bracketed word read aloud or skipped; with v3 it should become a pause. That single experiment teaches you more about "delivery cues" than any spec sheet.

Second worked example: a UK training company

A compliance-training firm in Leeds produces forty short modules a year. It picks Multilingual v2 for the narrator because consistency across hours of audio matters more than drama, and keeps v3 for the two-minute scenario dramatizations at the start of each module, where emotion helps learners remember. Both choices, with settings, sit in the voice spec, so a new producer can reproduce the sound months later.

Pitfalls

  • Choosing a voice from a single demo sentence. Test full scripts.
  • Changing several settings at once, so you cannot tell what helped.
  • Assuming multilingual means equally good in every language.
  • Forgetting licensing: check the platform's commercial-use terms for your plan.

Key takeaways

  • Modern neural TTS generates speech directly and can sound very natural.
  • Voice platforms offer TTS, speech-to-speech, voice design, cloning and dubbing.
  • Library, designed and cloned voices carry different uniqueness and consent implications.
  • Your script's punctuation, sentence length and spelling are your direction to the AI voice.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What does speech-to-speech (voice changing) let you do?
  2. A voice sounds flat and monotone. Which setting change is most likely to add expressiveness?
  3. Why create a pronunciation list for a brand?

Put it into practice

Test three voices on the same 60-second script in your main language, rate them on naturalness and pronunciation, and write a one-page voice spec for the winner.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.