AI Video & Voice Production: ElevenLabs, Veo, Runway and MoreAI voice: text-to-speech, cloning and dubbing · Lesson 4 of 16

Automating voice production with the ElevenLabs API

Article · 8 min · 8 min lecture

Video lecture

Automating voice production with the ElevenLabs API

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

Automating voice production

  • Script to scene audio
  • Skip unchanged scenes
  • One pronunciation list
  • A production log and human gates

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why automate voice production

Once a voice format works, the bottleneck is no longer quality. It is repetition: forty lessons, twelve product videos, a weekly show in three languages. Clicking through a web app for each line invites mistakes. Scripts do not get tired. This lesson shows how to turn a scene-based script into audio files, captions and a production log with the ElevenLabs Python SDK, and how to design the pipeline so a human still approves everything that ships.

You do not need to be a developer. You need to read a short script, change a few settings and run it. If you prefer no-code, the same logic works in automation tools that call the ElevenLabs API, but the concepts below stay the same.

The building blocks

JobSDK call (Python)Notes
Speak a lineclient.text_to_speech.convert(...)Choose voice_id, model_id, output_format
Stream for previewsclient.text_to_speech.stream(...)Useful for quick listening in notebooks
List voicesclient.voices.search()Find the ids of voices in your library
Design a voiceclient.text_to_voice.design(...)Previews from a description
Transcribe for captionsclient.speech_to_text.convert(...)Speech-to-text model (Scribe) with word timestamps
Dub a videoclient.dubbing.create(...)Async job; poll with client.dubbing.get(...)

Method names are taken from the SDK reference as of September 2026. SDKs evolve, so pin the version in requirements.txt and read the changelog before upgrading.

Design the pipeline before the code

A reliable pipeline has five properties:

  1. Scene-level files. One audio file per scene, named by scene number, so a single fix regenerates one file, not the whole lecture.
  2. Idempotent runs. If a file already exists and its text has not changed, skip it. Store a hash of the text next to the audio.
  3. Pronunciation in one place. Keep a dictionary of terms and their spoken forms and apply it before sending text.
  4. A production log. Record voice id, model id, settings, text hash, date and who approved it.
  5. Human gates. Nothing is published until a person has listened. Automation removes clicking, not judgment.

Hands-on: script to scene audio

Put your script in a CSV with columns scene,text. Then run:

# pip install elevenlabs==<pinned> python-dotenv
import csv, hashlib, json, os, pathlib
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs

load_dotenv()
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

VOICE_ID = os.environ["VOICE_ID"]          # from client.voices.search()
MODEL_ID = "eleven_multilingual_v2"         # stable long-form narration
OUT = pathlib.Path("audio"); OUT.mkdir(exist_ok=True)
SAY = {"GA4": "G A four", "SQL": "sequel", "KSA": "K S A"}   # pronunciation list

def spoken(text: str) -> str:
    for term, say in SAY.items():
        text = text.replace(term, say)
    return text

log = []
with open("script.csv", newline="", encoding="utf-8") as f:
    for row in csv.DictReader(f):
        text = spoken(row["text"].strip())
        digest = hashlib.sha256(f"{VOICE_ID}|{MODEL_ID}|{text}".encode()).hexdigest()[:12]
        path = OUT / f"scene_{int(row['scene']):02d}_{digest}.mp3"
        if path.exists():
            continue                                   # unchanged scene: skip
        try:
            audio = client.text_to_speech.convert(
                text=text, voice_id=VOICE_ID, model_id=MODEL_ID,
                output_format="mp3_44100_128")
            with open(path, "wb") as out:
                for chunk in audio:
                    out.write(chunk)
            log.append({"scene": row["scene"], "file": path.name, "hash": digest,
                        "voice": VOICE_ID, "model": MODEL_ID, "approved_by": None})
        except Exception as err:
            print(f"scene {row['scene']} failed: {err}")   # retry later; do not crash the batch

pathlib.Path("production_log.json").write_text(json.dumps(log, indent=2))

Two design choices matter. The hash in the file name means an edited line produces a new file while untouched scenes are skipped. And failures are logged rather than stopping the batch, so one bad line does not cost you the other thirty.

Captions from the final audio

Captions should match what was actually said, so transcribe the final mix rather than trusting the script:

with open("final_mix.mp3", "rb") as f:
    result = client.speech_to_text.convert(
        file=f, model_id="scribe_v2",            # check current STT model ids
        timestamps_granularity="word")
print(result.text[:200])

Use the word timestamps to build an SRT file or import the transcript into your editor. Then correct names and numbers by hand.

Worked example: a course team in Dubai

A training company produces eight lectures a month, each with ten scenes. Before automation, a producer spent most of a day per lecture clicking, downloading and renaming files, and scenes were sometimes regenerated with the wrong model. After building the script above, the producer edits the CSV, runs one command, listens to every scene at 1.25x speed, and marks approvals in the log. When a subject-matter expert changes one sentence, only one file regenerates. The team also discovered, through the log, that two lectures had been voiced with an older model; they re-rendered those for consistency.

Security and cost hygiene

  • Keep the API key in environment variables or a secrets manager; never commit .env files.
  • Use a separate API key per project so you can revoke one without breaking others, and restrict key permissions where the dashboard allows it.
  • Watch character usage in the dashboard. Idempotent runs keep costs predictable.
  • If your data is sensitive (unreleased products, client names), check the vendor's data-retention settings and your contract before sending text.

Pitfalls

  • One giant audio file per lecture, so a typo means re-rendering everything.
  • Unpinned SDK versions that change method names mid-project.
  • Automation without a listening pass. Every scene needs human ears before it ships.

Key takeaways

  • Automate repetition, not judgment: every scene still gets a human listening pass.
  • Generate one file per scene and fingerprint the text so unchanged scenes are skipped.
  • Keep pronunciations, voice ids and model ids in one place and log them for every file.
  • Transcribe the final mix for captions, pin SDK versions and keep API keys in environment variables.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why does the lesson include a hash of the voice, model and text in each audio file name?
  2. A producer wants captions that match exactly what the narrator says in the final video. What is the best source?
  3. Which practice best protects an automated voice pipeline from breaking unexpectedly?

Put it into practice

Put one real script into a scene,text CSV, run the lesson pipeline with a library voice, edit one line and confirm only that scene regenerates.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.