---
title: "How AI voice works: text-to-speech in plain language"
description: "From robotic to realistic Text-to-speech (TTS) converts written text into spoken audio. Earlier systems stitched together recorded sound fragments, which…"
url: https://optimizeall.com/learn/ai-video-and-voice-production/how-ai-voice-works
updated: 2026-10-05
---

AI Video & Voice Production: ElevenLabs, Veo, Runway and More · AI voice: text-to-speech, cloning and dubbing · lesson 1 of 16 · 9 min

# How AI voice works: text-to-speech in plain language

## From robotic to realistic

Text-to-speech (TTS) converts written text into spoken audio. Earlier systems stitched together recorded sound fragments, which is why they sounded choppy. Modern neural TTS models are trained on large amounts of recorded speech and generate audio directly, learning pronunciation, intonation, pacing, breath and emotional color. The best current voices are hard to distinguish from human narration in short clips.

Leading voice platforms such as ElevenLabs typically offer a family of related capabilities:

- **Text-to-speech:** type or paste a script, pick a voice, generate audio.
- **Speech-to-speech (voice changing):** record yourself with the timing and emotion you want, then convert it into another voice while keeping your performance.
- **Voice design:** create a new synthetic voice from a text description.
- **Voice cloning:** create a voice that replicates a specific real person, with that person's consent.
- **Dubbing and translation:** translate audio or video into other languages while preserving voice characteristics.
- **Transcription (speech-to-text), sound effects and conversational voice agents** as related tools.

Features, model names and plans change frequently, so check the current documentation of whichever tool you use. The section "The 2026 model line-up" below shows how to read a model menu.

## Key concepts

**Voice types**

- **Library or stock voices:** pre-made voices you can use under the platform's license. Fast, low risk, but other brands may use the same voice.
- **Designed voices:** generated from a description ("calm, deep, middle-aged male narrator with a neutral Gulf Arabic accent"). Unique to you, no real person's likeness involved.
- **Cloned voices:** trained on recordings of a real person. Best for creators who want to scale their own voice, and only with explicit consent.

**Models**

Platforms usually offer several models that trade off quality, emotional range, latency and language coverage. A fast, low-latency model suits live agents; a more expressive model suits ads and storytelling. Multilingual models can speak many languages, but quality varies by language and accent. Always test in your target language, for example Urdu, Arabic or Hindi, before committing.

**Settings you will meet**

Exact names vary, but common controls include:

- **Stability:** higher values give a more consistent, even delivery; lower values allow more variation and emotion, but also more unpredictability.
- **Similarity or clarity:** how closely output sticks to the original voice's characteristics.
- **Style or expressiveness:** exaggerates the voice's style; too much can sound theatrical.
- **Speed:** delivery pace.

Start with default settings, change one at a time and compare.

## The script is your direction

An AI narrator cannot ask what you meant. Your text is the direction:

- **Punctuation shapes pacing.** Commas create short pauses; full stops create longer ones; dashes and ellipses add hesitation.
- **Sentence length shapes energy.** Short sentences feel punchy. Long ones feel calm and explanatory.
- **Spell out tricky items.** Write "twenty twenty-six" or "2026" consistently, and spell out brand names, acronyms and local names phonetically if mispronounced. Many tools support pronunciation dictionaries or phonetic hints for recurring terms.
- **Some models accept delivery cues** in text, such as emotional or tonal tags, while others read everything literally. Test before relying on them.

## Worked example

Nadia, a Karachi-based e-commerce marketer, needs English and Urdu voiceovers for 20 product videos a month. She tests three library voices and one designed voice across five sample scripts in both languages, rating naturalness, pronunciation of product names and consistency. The designed voice wins in English; a different library voice handles Urdu better. She creates a pronunciation list for the brand name and product terms, and documents the chosen voice, model and settings in a one-page "voice spec" so every video sounds consistent.

## The 2026 model line-up (how to read it)

As of September 2026, ElevenLabs documents a small family of text-to-speech models, and the pattern generalizes to other vendors:

| Model (API id) | Best for | Trade-off |
|---|---|---|
| Eleven v3 (`eleven_v3`) | Expressive performances, ads, storytelling, multi-speaker dialogue; 70+ languages | Most emotional range; less predictable, so generate a few takes |
| Eleven Multilingual v2 (`eleven_multilingual_v2`) | Long-form narration, courses, audiobooks; stable delivery | Less dramatic range than v3 |
| Eleven Flash v2.5 (`eleven_flash_v2_5`) | Real-time agents and previews; very low latency | Slightly lower quality than the flagship models |

Eleven v3 also understands inline **audio tags** in square brackets, such as `[whispers]`, `[excited]` or `[sighs]`, which act as stage directions. Older models read them literally, which is why you always test a cue before you rely on it. Model names change: before a project, open the vendor's models page and write the model id into your voice spec.

## Hands-on: generate one line two ways

This short Python script uses the official `elevenlabs` SDK to render the same line with two models so you can compare them. Store your key in an environment variable, never in code.

```python
# pip install elevenlabs python-dotenv
import os
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs

load_dotenv()                      # reads ELEVENLABS_API_KEY from .env
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

VOICE_ID = os.environ.get("VOICE_ID", "JBFqnCBsd6RMkjVDRZzb")  # any voice from your library
LINE = "Your first ten orders are on us. [pause] Seriously, that's the offer."

for model in ["eleven_multilingual_v2", "eleven_v3"]:
    try:
        audio = client.text_to_speech.convert(
            text=LINE, voice_id=VOICE_ID, model_id=model,
            output_format="mp3_44100_128",
        )
        with open(f"take_{model}.mp3", "wb") as f:
            for chunk in audio:          # the SDK returns the audio as a stream of bytes
                f.write(chunk)
        print("saved", model)
    except Exception as err:             # quota, invalid voice id, network
        print("failed", model, err)
```

Listen to both takes. With Multilingual v2 you will probably hear the bracketed word read aloud or skipped; with v3 it should become a pause. That single experiment teaches you more about "delivery cues" than any spec sheet.

## Second worked example: a UK training company

A compliance-training firm in Leeds produces forty short modules a year. It picks Multilingual v2 for the narrator because consistency across hours of audio matters more than drama, and keeps v3 for the two-minute scenario dramatizations at the start of each module, where emotion helps learners remember. Both choices, with settings, sit in the voice spec, so a new producer can reproduce the sound months later.

## Pitfalls

- Choosing a voice from a single demo sentence. Test full scripts.
- Changing several settings at once, so you cannot tell what helped.
- Assuming multilingual means equally good in every language.
- Forgetting licensing: check the platform's commercial-use terms for your plan.

## Video lecture: How AI voice works: text-to-speech in plain language

Lecture coming soon · 13 chapters · about 8 minutes. Read the full transcript below.

1. How AI voice works
2. Why it matters
3. The core idea
4. Reading the model menu
5. The four dials
6. Example 1: a 30-second teaser
7. Example 2: Karachi e-commerce
8. Watch me do it
9. The voice spec
10. Common mistakes
11. Measuring success
12. Recap
13. Try this now

## Lecture transcript

### How AI voice works

Close your eyes for a second and imagine the last ad you heard on a podcast. Could you tell whether that voice was a person or a model? In twenty twenty-six, often you can't. In this lecture you'll learn how modern AI voices actually work, how to read a vendor's model menu, and how to pick and test a voice so that every video you make sounds consistent. By the end, you'll be able to write a one-page voice spec that anyone on your team can follow.

### Why it matters

Why does this matter? Because the voice is the part of a video people forgive least. Viewers will tolerate simple visuals, but a voice that mispronounces your brand, drifts in tone halfway through, or sounds flat will make them scroll. And for creators and small teams, voice is also where AI saves the most time. A voiceover that used to need a booth, a narrator and a retake session can now be generated, corrected and re-rendered in minutes. The skill is no longer recording. It's directing.

### The core idea

Here's the core idea. Old text-to-speech glued together tiny recorded fragments, which is why it sounded choppy. Modern neural models learned from huge amounts of recorded speech, so they generate audio directly, including breath, rhythm and emotion. Think of it like a session musician who has heard thousands of songs. You don't hand them individual notes. You hand them a chart and a feel, and they perform. Your script is the chart. Punctuation, sentence length and a few cues are the feel. The model fills in the rest.

### Reading the model menu

Now let's read a real model menu. As of September twenty twenty-six, ElevenLabs lists three main text-to-speech models. Eleven v3 is the expressive one, great for ads, stories and dialogue, and it understands little stage directions called audio tags, like whispers or excited. Multilingual v2 is the steady one, best for long narration where every paragraph should sound the same. Flash two point five is the fast one, built for live agents where speed beats polish. The names will change. The pattern won't: expressive, stable, fast. Pick for the job.

### The four dials

Next, the settings you'll meet. Exact names vary, but most tools give you four dials. Stability controls how even the delivery is. Higher means consistent, lower means more emotion and more surprises. Similarity controls how closely the output sticks to the original voice. Style exaggerates the voice's character, and too much sounds theatrical. And speed is pace. Here's the discipline that separates pros from hobbyists. Start with the defaults. Change one dial at a time. Render the same paragraph after each change and compare them side by side. Otherwise you'll never know which change actually helped.

### Example 1: a 30-second teaser

First example, a simple one. You need a thirty-second product teaser. You paste the script into the voice tool, choose a library voice, and generate. It sounds fine, except the brand name comes out wrong and the last line feels rushed. Two fixes. You add the brand to a pronunciation list, spelled the way it sounds. Then you break the last sentence in two and add a full stop, which gives the voice a beat to breathe. Generate again. Same voice, same model, and now it lands. Notice you changed the text, not the settings.

### Example 2: Karachi e-commerce

Second example, a realistic business. Illustrative numbers. A Karachi e-commerce team needs about twenty English and Urdu voiceovers a month. Their marketer, Nadia, doesn't pick from a single demo sentence. She takes five real scripts and renders them with three library voices and one designed voice, in both languages. She scores each on naturalness, pronunciation of product names, and consistency across scripts. The designed voice wins in English. A different library voice wins in Urdu. She writes both into a voice spec with the model id and settings. Now anyone on the team can reproduce the brand sound.

### Watch me do it

Watch me do it. I open a small Python script from the lesson. It loads my key from an environment variable, never from the code. I set a voice id and one test line that includes a pause cue in square brackets. Then I loop over two models, Multilingual v2 and Eleven v3, and save each take as its own file. I run it. Two files appear. I play the first. The bracketed cue gets read out or skipped. I play the second. It becomes a real pause. That one comparison tells me which model supports cues, and I write it down in the spec.

### The voice spec

Now I turn that test into a voice spec. It's one page, and it has six lines. The voice name and id. The model id, because model names change. The settings, exactly as used. The pronunciation list, with each brand or local name spelled the way it should sound. The output format. And one reference file, a thirty-second approved render that anyone can compare against. I save it next to the project. Next month, when a freelancer joins, they don't guess. They open the spec, render a test, and compare it with the reference.

### Common mistakes

Now the common mistakes. One, choosing a voice from a single demo sentence. Demos are cherry-picked, so always test full scripts. Two, changing several settings at once, so you can't tell what helped. Change one, compare, then move on. Three, assuming multilingual means equally good in every language. Quality varies by language and accent, so test in Urdu, Arabic or Hindi before promising a client. Four, forgetting the license. Check that your plan allows commercial use of the audio before it goes into an ad.

### Measuring success

How do you know it's working? Measure three things. First, rework rate: how many scenes you had to regenerate per video. It should fall as your spec and pronunciation list mature. Second, consistency: play the first and last video of the month back to back. Do they sound like the same narrator? Third, audience signals: completion rate and comments about the voice. If people mention the voice, good or bad, that's data. A falling rework rate and steady completion rates tell you your voice system is healthy.

### Recap

Quick recap. Modern voices generate a performance rather than stitching fragments. Model menus follow a pattern: expressive, stable and fast, so pick for the job. Your script is the direction, so fix the text before the settings. Test with real scripts in every language you need, score them, and write the winner into a voice spec with the model id and settings. Then keep an eye on your rework rate. If you're still regenerating half your scenes after a month, the problem is usually the script or the pronunciation list, not the voice.

### Try this now

Try this now. Take one sixty-second script you actually need this month. Render it with three voices in your main language. Score each out of five for naturalness, pronunciation and consistency. Then open a blank page and write your voice spec: voice, model id, settings, and your pronunciation list. It takes twenty minutes, and it will save you hours of re-rendering later.

## Video transcript

In this course, we're going to produce professional voice and video content with AI. Let's start with voice, because it underpins almost everything else.

Text-to-speech, or TTS, turns written text into spoken audio. Older systems sounded robotic because they stitched together recorded fragments. Modern AI voices, from tools like ElevenLabs, are generated by models trained on large amounts of speech. They learn not just pronunciation, but rhythm, emphasis, breath and emotion. That's why today's best voices can sound remarkably human.

You'll work with three kinds of voices. Library voices, which are ready-made voices licensed for use under the tool's terms. Designed voices, where you describe the voice you want, such as a warm, mid-thirties female narrator with a light British accent, and the tool generates it. And cloned voices, which replicate a real person's voice from recordings. Cloning is powerful, and it comes with serious consent and rights responsibilities, which we'll cover next.

The quality of your output depends on three things. The model you choose, because some prioritize expressiveness and others speed. Your settings, like stability and similarity, which trade consistency against emotional range. And most importantly, your script. AI voices read exactly what you write, so punctuation, spelling and structure become your direction to the performer.

By the end of this module, you'll know how to choose, design and direct AI voices, and how to use them to reach audiences in new languages, legally and ethically.

## Key takeaways

- Modern neural TTS generates speech directly and can sound very natural.
- Voice platforms offer TTS, speech-to-speech, voice design, cloning and dubbing.
- Library, designed and cloned voices carry different uniqueness and consent implications.
- Your script's punctuation, sentence length and spelling are your direction to the AI voice.

## Try it

Test three voices on the same 60-second script in your main language, rate them on naturalness and pronunciation, and write a one-page voice spec for the winner.

- [Next: Voice design and cloning: consent and rights first](https://optimizeall.com/learn/ai-video-and-voice-production/voice-design-and-cloning-with-consent)
- [All lessons of AI Video & Voice Production: ElevenLabs, Veo, Runway and More](https://optimizeall.com/learn/ai-video-and-voice-production)
