AI Video & Voice Production: ElevenLabs, Veo, Runway and MoreAI voice: text-to-speech, cloning and dubbing · Lesson 1 of 16
How AI voice works: text-to-speech in plain language
Video lecture
How AI voice works: text-to-speech in plain language
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 How AI voice works
Close your eyes for a second and imagine the last ad you heard on a podcast. Could you tell whether that voice was a person or a model? In twenty twenty-six, often you can't. In this lecture you'll learn how modern AI voices actually work, how to read a vendor's model menu, and how to pick and test a voice so that every video you make sounds consistent. By the end, you'll be able to write a one-page voice spec that anyone on your team can follow.
0:38 Why it matters
Why does this matter? Because the voice is the part of a video people forgive least. Viewers will tolerate simple visuals, but a voice that mispronounces your brand, drifts in tone halfway through, or sounds flat will make them scroll. And for creators and small teams, voice is also where AI saves the most time. A voiceover that used to need a booth, a narrator and a retake session can now be generated, corrected and re-rendered in minutes. The skill is no longer recording. It's directing.
1:15 The core idea
Here's the core idea. Old text-to-speech glued together tiny recorded fragments, which is why it sounded choppy. Modern neural models learned from huge amounts of recorded speech, so they generate audio directly, including breath, rhythm and emotion. Think of it like a session musician who has heard thousands of songs. You don't hand them individual notes. You hand them a chart and a feel, and they perform. Your script is the chart. Punctuation, sentence length and a few cues are the feel. The model fills in the rest.
1:53 Reading the model menu
Now let's read a real model menu. As of September twenty twenty-six, ElevenLabs lists three main text-to-speech models. Eleven v3 is the expressive one, great for ads, stories and dialogue, and it understands little stage directions called audio tags, like whispers or excited. Multilingual v2 is the steady one, best for long narration where every paragraph should sound the same. Flash two point five is the fast one, built for live agents where speed beats polish. The names will change. The pattern won't: expressive, stable, fast. Pick for the job.
2:32 The four dials
Next, the settings you'll meet. Exact names vary, but most tools give you four dials. Stability controls how even the delivery is. Higher means consistent, lower means more emotion and more surprises. Similarity controls how closely the output sticks to the original voice. Style exaggerates the voice's character, and too much sounds theatrical. And speed is pace. Here's the discipline that separates pros from hobbyists. Start with the defaults. Change one dial at a time. Render the same paragraph after each change and compare them side by side. Otherwise you'll never know which change actually helped.
3:14 Example 1: a 30-second teaser
First example, a simple one. You need a thirty-second product teaser. You paste the script into the voice tool, choose a library voice, and generate. It sounds fine, except the brand name comes out wrong and the last line feels rushed. Two fixes. You add the brand to a pronunciation list, spelled the way it sounds. Then you break the last sentence in two and add a full stop, which gives the voice a beat to breathe. Generate again. Same voice, same model, and now it lands. Notice you changed the text, not the settings.
3:55 Example 2: Karachi e-commerce
Second example, a realistic business. Illustrative numbers. A Karachi e-commerce team needs about twenty English and Urdu voiceovers a month. Their marketer, Nadia, doesn't pick from a single demo sentence. She takes five real scripts and renders them with three library voices and one designed voice, in both languages. She scores each on naturalness, pronunciation of product names, and consistency across scripts. The designed voice wins in English. A different library voice wins in Urdu. She writes both into a voice spec with the model id and settings. Now anyone on the team can reproduce the brand sound.
4:38 Watch me do it
Watch me do it. I open a small Python script from the lesson. It loads my key from an environment variable, never from the code. I set a voice id and one test line that includes a pause cue in square brackets. Then I loop over two models, Multilingual v2 and Eleven v3, and save each take as its own file. I run it. Two files appear. I play the first. The bracketed cue gets read out or skipped. I play the second. It becomes a real pause. That one comparison tells me which model supports cues, and I write it down in the spec.
5:24 The voice spec
Now I turn that test into a voice spec. It's one page, and it has six lines. The voice name and id. The model id, because model names change. The settings, exactly as used. The pronunciation list, with each brand or local name spelled the way it should sound. The output format. And one reference file, a thirty-second approved render that anyone can compare against. I save it next to the project. Next month, when a freelancer joins, they don't guess. They open the spec, render a test, and compare it with the reference.
6:05 Common mistakes
Now the common mistakes. One, choosing a voice from a single demo sentence. Demos are cherry-picked, so always test full scripts. Two, changing several settings at once, so you can't tell what helped. Change one, compare, then move on. Three, assuming multilingual means equally good in every language. Quality varies by language and accent, so test in Urdu, Arabic or Hindi before promising a client. Four, forgetting the license. Check that your plan allows commercial use of the audio before it goes into an ad.
6:42 Measuring success
How do you know it's working? Measure three things. First, rework rate: how many scenes you had to regenerate per video. It should fall as your spec and pronunciation list mature. Second, consistency: play the first and last video of the month back to back. Do they sound like the same narrator? Third, audience signals: completion rate and comments about the voice. If people mention the voice, good or bad, that's data. A falling rework rate and steady completion rates tell you your voice system is healthy.
7:20 Recap
Quick recap. Modern voices generate a performance rather than stitching fragments. Model menus follow a pattern: expressive, stable and fast, so pick for the job. Your script is the direction, so fix the text before the settings. Test with real scripts in every language you need, score them, and write the winner into a voice spec with the model id and settings. Then keep an eye on your rework rate. If you're still regenerating half your scenes after a month, the problem is usually the script or the pronunciation list, not the voice.
8:00 Try this now
Try this now. Take one sixty-second script you actually need this month. Render it with three voices in your main language. Score each out of five for naturalness, pronunciation and consistency. Then open a blank page and write your voice spec: voice, model id, settings, and your pronunciation list. It takes twenty minutes, and it will save you hours of re-rendering later.
From robotic to realistic
Text-to-speech (TTS) converts written text into spoken audio. Earlier systems stitched together recorded sound fragments, which is why they sounded choppy. Modern neural TTS models are trained on large amounts of recorded speech and generate audio directly, learning pronunciation, intonation, pacing, breath and emotional color. The best current voices are hard to distinguish from human narration in short clips.
Leading voice platforms such as ElevenLabs typically offer a family of related capabilities:
- Text-to-speech: type or paste a script, pick a voice, generate audio.
- Speech-to-speech (voice changing): record yourself with the timing and emotion you want, then convert it into another voice while keeping your performance.
- Voice design: create a new synthetic voice from a text description.
- Voice cloning: create a voice that replicates a specific real person, with that person's consent.
- Dubbing and translation: translate audio or video into other languages while preserving voice characteristics.
- Transcription (speech-to-text), sound effects and conversational voice agents as related tools.
Features, model names and plans change frequently, so check the current documentation of whichever tool you use. The section "The 2026 model line-up" below shows how to read a model menu.
Key concepts
Voice types
- Library or stock voices: pre-made voices you can use under the platform's license. Fast, low risk, but other brands may use the same voice.
- Designed voices: generated from a description ("calm, deep, middle-aged male narrator with a neutral Gulf Arabic accent"). Unique to you, no real person's likeness involved.
- Cloned voices: trained on recordings of a real person. Best for creators who want to scale their own voice, and only with explicit consent.
Models
Platforms usually offer several models that trade off quality, emotional range, latency and language coverage. A fast, low-latency model suits live agents; a more expressive model suits ads and storytelling. Multilingual models can speak many languages, but quality varies by language and accent. Always test in your target language, for example Urdu, Arabic or Hindi, before committing.
Settings you will meet
Exact names vary, but common controls include:
- Stability: higher values give a more consistent, even delivery; lower values allow more variation and emotion, but also more unpredictability.
- Similarity or clarity: how closely output sticks to the original voice's characteristics.
- Style or expressiveness: exaggerates the voice's style; too much can sound theatrical.
- Speed: delivery pace.
Start with default settings, change one at a time and compare.
The script is your direction
An AI narrator cannot ask what you meant. Your text is the direction:
- Punctuation shapes pacing. Commas create short pauses; full stops create longer ones; dashes and ellipses add hesitation.
- Sentence length shapes energy. Short sentences feel punchy. Long ones feel calm and explanatory.
- Spell out tricky items. Write "twenty twenty-six" or "2026" consistently, and spell out brand names, acronyms and local names phonetically if mispronounced. Many tools support pronunciation dictionaries or phonetic hints for recurring terms.
- Some models accept delivery cues in text, such as emotional or tonal tags, while others read everything literally. Test before relying on them.
Worked example
Nadia, a Karachi-based e-commerce marketer, needs English and Urdu voiceovers for 20 product videos a month. She tests three library voices and one designed voice across five sample scripts in both languages, rating naturalness, pronunciation of product names and consistency. The designed voice wins in English; a different library voice handles Urdu better. She creates a pronunciation list for the brand name and product terms, and documents the chosen voice, model and settings in a one-page "voice spec" so every video sounds consistent.
The 2026 model line-up (how to read it)
As of September 2026, ElevenLabs documents a small family of text-to-speech models, and the pattern generalizes to other vendors:
| Model (API id) | Best for | Trade-off |
|---|---|---|
Eleven v3 (eleven_v3) | Expressive performances, ads, storytelling, multi-speaker dialogue; 70+ languages | Most emotional range; less predictable, so generate a few takes |
Eleven Multilingual v2 (eleven_multilingual_v2) | Long-form narration, courses, audiobooks; stable delivery | Less dramatic range than v3 |
Eleven Flash v2.5 (eleven_flash_v2_5) | Real-time agents and previews; very low latency | Slightly lower quality than the flagship models |
Eleven v3 also understands inline audio tags in square brackets, such as [whispers], [excited] or [sighs], which act as stage directions. Older models read them literally, which is why you always test a cue before you rely on it. Model names change: before a project, open the vendor's models page and write the model id into your voice spec.
Hands-on: generate one line two ways
This short Python script uses the official elevenlabs SDK to render the same line with two models so you can compare them. Store your key in an environment variable, never in code.
# pip install elevenlabs python-dotenv
import os
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv() # reads ELEVENLABS_API_KEY from .env
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
VOICE_ID = os.environ.get("VOICE_ID", "JBFqnCBsd6RMkjVDRZzb") # any voice from your library
LINE = "Your first ten orders are on us. [pause] Seriously, that's the offer."
for model in ["eleven_multilingual_v2", "eleven_v3"]:
try:
audio = client.text_to_speech.convert(
text=LINE, voice_id=VOICE_ID, model_id=model,
output_format="mp3_44100_128",
)
with open(f"take_{model}.mp3", "wb") as f:
for chunk in audio: # the SDK returns the audio as a stream of bytes
f.write(chunk)
print("saved", model)
except Exception as err: # quota, invalid voice id, network
print("failed", model, err)Listen to both takes. With Multilingual v2 you will probably hear the bracketed word read aloud or skipped; with v3 it should become a pause. That single experiment teaches you more about "delivery cues" than any spec sheet.
Second worked example: a UK training company
A compliance-training firm in Leeds produces forty short modules a year. It picks Multilingual v2 for the narrator because consistency across hours of audio matters more than drama, and keeps v3 for the two-minute scenario dramatizations at the start of each module, where emotion helps learners remember. Both choices, with settings, sit in the voice spec, so a new producer can reproduce the sound months later.
Pitfalls
- Choosing a voice from a single demo sentence. Test full scripts.
- Changing several settings at once, so you cannot tell what helped.
- Assuming multilingual means equally good in every language.
- Forgetting licensing: check the platform's commercial-use terms for your plan.
Key takeaways
- Modern neural TTS generates speech directly and can sound very natural.
- Voice platforms offer TTS, speech-to-speech, voice design, cloning and dubbing.
- Library, designed and cloned voices carry different uniqueness and consent implications.
- Your script's punctuation, sentence length and spelling are your direction to the AI voice.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Test three voices on the same 60-second script in your main language, rate them on naturalness and pronunciation, and write a one-page voice spec for the winner.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.