AI Video & Voice Production: ElevenLabs, Veo, Runway and MoreAI voice: text-to-speech, cloning and dubbing · Lesson 4 of 16
Automating voice production with the ElevenLabs API
Video lecture
Automating voice production with the ElevenLabs API
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Automating voice production
If you've ever re-rendered a whole ten-minute voiceover because of one typo, this lecture is for you. You'll learn to turn a scene-based script into audio files with a short Python script, skip scenes that haven't changed, keep pronunciation in one place, and log exactly which voice and model produced every file. And you'll keep a human in the loop, because automation should remove clicking, not judgment.
0:29 Why automate
Why automate? Because once a voice format works, quality stops being the bottleneck. Repetition is. Forty lessons, twelve product videos, a weekly show in three languages. Every manual click is a chance to pick the wrong voice, the wrong model or last week's text. Scripts don't get tired and they don't forget settings. But the goal isn't a robot factory. It's a system where the boring steps are automatic and the judgment steps are deliberate.
1:02 The kitchen model
Here's the mental model. Think of your pipeline like a professional kitchen. Recipes are the script. Mise en place is the pronunciation list and voice settings prepared once. Each dish is a scene, cooked separately so one burnt plate doesn't mean remaking the whole meal. The ticket rail is the production log. And the head chef tastes everything before it leaves the pass. That's your human listening check.
1:32 Building blocks
The building blocks are few. Text to speech convert turns one line into audio with a voice id, a model id and an output format. Voices search lists the ids in your library. Speech to text convert transcribes audio for captions, with word timestamps. And dubbing create starts a translation job you check on later. Those names come from the official SDK reference as of September twenty twenty-six. Pin the SDK version in your project, and read the changelog before you upgrade.
2:08 Five pipeline properties
Before any code, design five properties. Scene-level files, so a fix touches one file. Idempotent runs, which means running twice produces the same result and skips finished work. Pronunciation in one place. A production log recording voice, model, settings, a fingerprint of the text, the date and who approved it. And human gates, meaning nothing ships until a person has listened. If a pipeline lacks any of these five, it will eventually waste money, produce inconsistent audio, or publish something nobody checked.
2:44 Example 1: only what changed
First example, the simplest possible win. You have a ten-scene lecture in a spreadsheet with two columns, scene and text. The script reads each row, sends the text to the voice, and saves scene zero one, scene zero two and so on. Your expert changes one sentence in scene seven. Instead of re-rendering everything, you rerun the script and only scene seven is regenerated. That's because each file name includes a short fingerprint of its text. Same text, same fingerprint, skip.
3:19 Example 2: Dubai training team
Second example, a business case with illustrative numbers. A Dubai training company produces eight lectures a month with about ten scenes each. Before automation, a producer spent most of a day per lecture clicking, downloading and renaming, and occasionally used the wrong model. After the script, the producer edits the spreadsheet, runs one command, and listens to every scene at one and a quarter speed, marking approvals in the log. The log also exposed that two older lectures used a different model, so they re-rendered those for consistency. Time moved from clicking to listening.
4:00 Watch me do it
Watch me do it. I open the lesson's script. At the top, it loads my API key from the environment, and the voice id too. I set the model to Multilingual v2 for stable narration. Next, a small dictionary: G A four, sequel, K S A. That's my pronunciation list, applied before any text is sent. Then the loop. For each row, it builds a fingerprint from the voice, the model and the text. If a file with that fingerprint exists, it skips. Otherwise it generates, saves, and appends a line to the log. If anything fails, it prints the error and keeps going.
4:45 Captions and sign-off
Now captions. After the final mix is edited, I transcribe it rather than trusting the script, because scripts drift during recording and editing. I call speech to text convert on the final file, ask for word-level timestamps, and get back text with timings. I turn that into a subtitle file or import it into the editor, then fix names and numbers by hand. Finally, I open the production log, find each scene I listened to, and fill in my name as approver. Only then does the lecture move to publishing.
5:24 Security and cost hygiene
And security, briefly, because it's where small teams get burned. Keep the key in an environment variable or a secrets manager, never in the code, and never commit your dot env file. Use a separate key per project, so you can revoke one without breaking everything. Watch your character usage in the dashboard; skipping unchanged scenes keeps costs predictable. And if your scripts mention unreleased products or client names, check the vendor's data retention settings and your contract before sending them.
5:59 Common mistakes
Now the common mistakes. One giant audio file per lecture, so every typo costs a full re-render. Unpinned SDK versions, so a method name changes mid-project and the script breaks the night before launch. API keys pasted into code and committed to a shared repository. And the big one, automation without a listening pass. A script will happily produce thirty scenes with the wrong brand pronunciation. Human ears are the last gate.
6:30 Recap
Recap. Generate audio per scene. Fingerprint the text so unchanged scenes are skipped. Keep pronunciations in one dictionary. Log the voice, model and approvals. Transcribe the final mix for captions. And never ship without a person listening. Keep keys out of code, one key per project, and pin your SDK version. The payoff is a pipeline where fixing a sentence takes seconds, and where you can prove exactly how every file was made.
7:02 Try this now
Try this now. Put one of your scripts into a two-column spreadsheet, scene and text. Run the lesson's script with a library voice. Then change one line, rerun, and confirm only that scene regenerates. If that works, you've built the core of a production pipeline. Next, add three terms to the pronunciation dictionary that your voice gets wrong, and rerun. Finally, open the log and check that every file shows the voice, the model and a blank approval field waiting for you. Listen, approve, and you're done.
Why automate voice production
Once a voice format works, the bottleneck is no longer quality. It is repetition: forty lessons, twelve product videos, a weekly show in three languages. Clicking through a web app for each line invites mistakes. Scripts do not get tired. This lesson shows how to turn a scene-based script into audio files, captions and a production log with the ElevenLabs Python SDK, and how to design the pipeline so a human still approves everything that ships.
You do not need to be a developer. You need to read a short script, change a few settings and run it. If you prefer no-code, the same logic works in automation tools that call the ElevenLabs API, but the concepts below stay the same.
The building blocks
| Job | SDK call (Python) | Notes |
|---|---|---|
| Speak a line | client.text_to_speech.convert(...) | Choose voice_id, model_id, output_format |
| Stream for previews | client.text_to_speech.stream(...) | Useful for quick listening in notebooks |
| List voices | client.voices.search() | Find the ids of voices in your library |
| Design a voice | client.text_to_voice.design(...) | Previews from a description |
| Transcribe for captions | client.speech_to_text.convert(...) | Speech-to-text model (Scribe) with word timestamps |
| Dub a video | client.dubbing.create(...) | Async job; poll with client.dubbing.get(...) |
Method names are taken from the SDK reference as of September 2026. SDKs evolve, so pin the version in requirements.txt and read the changelog before upgrading.
Design the pipeline before the code
A reliable pipeline has five properties:
- Scene-level files. One audio file per scene, named by scene number, so a single fix regenerates one file, not the whole lecture.
- Idempotent runs. If a file already exists and its text has not changed, skip it. Store a hash of the text next to the audio.
- Pronunciation in one place. Keep a dictionary of terms and their spoken forms and apply it before sending text.
- A production log. Record voice id, model id, settings, text hash, date and who approved it.
- Human gates. Nothing is published until a person has listened. Automation removes clicking, not judgment.
Hands-on: script to scene audio
Put your script in a CSV with columns scene,text. Then run:
# pip install elevenlabs==<pinned> python-dotenv
import csv, hashlib, json, os, pathlib
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv()
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
VOICE_ID = os.environ["VOICE_ID"] # from client.voices.search()
MODEL_ID = "eleven_multilingual_v2" # stable long-form narration
OUT = pathlib.Path("audio"); OUT.mkdir(exist_ok=True)
SAY = {"GA4": "G A four", "SQL": "sequel", "KSA": "K S A"} # pronunciation list
def spoken(text: str) -> str:
for term, say in SAY.items():
text = text.replace(term, say)
return text
log = []
with open("script.csv", newline="", encoding="utf-8") as f:
for row in csv.DictReader(f):
text = spoken(row["text"].strip())
digest = hashlib.sha256(f"{VOICE_ID}|{MODEL_ID}|{text}".encode()).hexdigest()[:12]
path = OUT / f"scene_{int(row['scene']):02d}_{digest}.mp3"
if path.exists():
continue # unchanged scene: skip
try:
audio = client.text_to_speech.convert(
text=text, voice_id=VOICE_ID, model_id=MODEL_ID,
output_format="mp3_44100_128")
with open(path, "wb") as out:
for chunk in audio:
out.write(chunk)
log.append({"scene": row["scene"], "file": path.name, "hash": digest,
"voice": VOICE_ID, "model": MODEL_ID, "approved_by": None})
except Exception as err:
print(f"scene {row['scene']} failed: {err}") # retry later; do not crash the batch
pathlib.Path("production_log.json").write_text(json.dumps(log, indent=2))Two design choices matter. The hash in the file name means an edited line produces a new file while untouched scenes are skipped. And failures are logged rather than stopping the batch, so one bad line does not cost you the other thirty.
Captions from the final audio
Captions should match what was actually said, so transcribe the final mix rather than trusting the script:
with open("final_mix.mp3", "rb") as f:
result = client.speech_to_text.convert(
file=f, model_id="scribe_v2", # check current STT model ids
timestamps_granularity="word")
print(result.text[:200])Use the word timestamps to build an SRT file or import the transcript into your editor. Then correct names and numbers by hand.
Worked example: a course team in Dubai
A training company produces eight lectures a month, each with ten scenes. Before automation, a producer spent most of a day per lecture clicking, downloading and renaming files, and scenes were sometimes regenerated with the wrong model. After building the script above, the producer edits the CSV, runs one command, listens to every scene at 1.25x speed, and marks approvals in the log. When a subject-matter expert changes one sentence, only one file regenerates. The team also discovered, through the log, that two lectures had been voiced with an older model; they re-rendered those for consistency.
Security and cost hygiene
- Keep the API key in environment variables or a secrets manager; never commit
.envfiles. - Use a separate API key per project so you can revoke one without breaking others, and restrict key permissions where the dashboard allows it.
- Watch character usage in the dashboard. Idempotent runs keep costs predictable.
- If your data is sensitive (unreleased products, client names), check the vendor's data-retention settings and your contract before sending text.
Pitfalls
- One giant audio file per lecture, so a typo means re-rendering everything.
- Unpinned SDK versions that change method names mid-project.
- Automation without a listening pass. Every scene needs human ears before it ships.
Key takeaways
- Automate repetition, not judgment: every scene still gets a human listening pass.
- Generate one file per scene and fingerprint the text so unchanged scenes are skipped.
- Keep pronunciations, voice ids and model ids in one place and log them for every file.
- Transcribe the final mix for captions, pin SDK versions and keep API keys in environment variables.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Put one real script into a scene,text CSV, run the lesson pipeline with a library voice, edit one line and confirm only that scene regenerates.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.