---
title: "Produced audio: narration, podcasts and course lectures"
description: "A different job from live agents Not all voice AI is conversational. Produced audio (course lectures, explainer videos, audiobooks, podcast segments…"
url: https://optimizeall.com/learn/voice-ai-agents/narration-podcasts-and-course-lectures
updated: 2026-10-05
---

Voice AI & Conversational Agents · Use cases and capstone · lesson 16 of 17 · 15 min

# Produced audio: narration, podcasts and course lectures

## A different job from live agents

Not all voice AI is conversational. **Produced audio** (course lectures, explainer videos, audiobooks, podcast segments, ads, IVR prompts, localized versions of existing content) is generated offline, reviewed, and published. Latency does not matter; quality, consistency, pronunciation and rights do. This very platform uses ElevenLabs voices to narrate course lectures over AI-generated visuals.

## The production pipeline

```text
Script (written for the ear) -> Pronunciation pass -> Voice + model selection
 -> Generate in segments -> Listen and fix -> Assemble with music/visuals
 -> Loudness normalization -> Captions/transcript -> Disclosure + metadata -> Publish
```

## Writing scripts for narration

- One idea per sentence; vary sentence length for rhythm.
- Hook in the first 15 seconds; state what the listener will be able to do.
- Signpost structure ("Three steps. First...").
- Spell out numbers, units and acronyms as they should be spoken, or use pronunciation rules.
- Avoid reading URLs; say "the link in the lesson notes".
- For educational content, use concrete examples and a recap with a next step.

## Model and voice choices for produced audio

- **Expressive long-form**: Eleven v3 with audio tags (`[excited]`, `[whispers]`, `[pause]`-style direction, where supported) for performance; test that tags behave as expected for your voice.
- **Consistency over many hours**: Multilingual v2 is known for stability across long content; lock voice settings and model versions for a series.
- **Multi-speaker**: dialogue features or separate voices per speaker for podcast-style segments.
- **Dubbing and localization**: ElevenLabs offers dubbing that translates and re-voices audio/video while trying to preserve speaker characteristics; always have native speakers review.

## Consistency across a course or series

- Fix voice ID, model, stability/similarity/speed settings and output format in a config file.
- Generate per scene or paragraph to make fixes cheap; keep a manifest mapping script segments to audio files.
- Use request stitching or context features (where available) so adjacent segments sound continuous.
- Keep a pronunciation dictionary for brand names, technical terms and people's names.

## Quality control

- Listen to everything before publishing. Check mispronunciations, odd emphasis, clipped endings, repeated words and artifacts.
- Normalize loudness (for example, podcasts commonly target around -16 LUFS integrated for stereo; check your platform's recommendation).
- Generate captions from the script or with speech-to-text (for example Scribe) and correct them; captions improve accessibility and SEO.

## Rights and disclosure

- Use voices you have rights to: library voices under their terms, designed voices, or consented clones.
- Disclose AI narration where required (EU AI Act for deepfake audio of real people; platform rules) and where your audience would reasonably expect to know. A simple line such as "Narrated with an AI voice" in the description is good practice.
- Podcast and audiobook platforms have their own AI-narration policies; check before uploading.

## Worked example: narrating a 17-lesson course

1. Scripts written per scene (5 to 16 scenes per lecture, 40 to 260 words each).
2. Voice: designed "calm expert" voice; model locked; settings in `voice_config.json`.
3. Pronunciation dictionary: GA4 -> "G A four", SIP -> "sip", brand names.
4. Batch generation per scene with retries; manifest records scene ID, text hash, audio file and duration.
5. QA listening pass; regenerate flagged scenes only.
6. Assemble with visuals; normalize loudness; captions from script; publish with "AI-narrated" note.

## Hands-on: batch-narrate scenes with a manifest

```python
import os, json, hashlib, time
from elevenlabs.client import ElevenLabs

client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
cfg = json.load(open("voice_config.json"))  # {"voice_id": "...", "model_id": "eleven_multilingual_v2", "output_format": "mp3_44100_128"}
scenes = json.load(open("lecture_scenes.json"))  # [{"id": "l01-s01", "narration": "..."}]
manifest = {}

for s in scenes:
    text_hash = hashlib.sha256(s["narration"].encode()).hexdigest()[:12]
    out = f"audio/{s['id']}-{text_hash}.mp3"
    if os.path.exists(out):
        manifest[s["id"]] = out  # unchanged text: reuse
        continue
    for attempt in range(3):
        try:
            audio = client.text_to_speech.convert(voice_id=cfg["voice_id"], model_id=cfg["model_id"],
                                                  text=s["narration"], output_format=cfg["output_format"])
            os.makedirs("audio", exist_ok=True)
            with open(out, "wb") as f:
                for chunk in audio:
                    f.write(chunk)
            manifest[s["id"]] = out
            break
        except Exception as e:
            print(f"{s['id']} attempt {attempt + 1} failed: {type(e).__name__}")
            time.sleep(2 ** attempt)

json.dump(manifest, open("manifest.json", "w"), indent=2)
```

Hashing the text means only edited scenes are regenerated, which saves cost and keeps approved audio stable.

## Pitfalls

- Changing voice settings mid-series, so lessons sound like different narrators.
- Skipping the listening pass.
- Using a voice without clear rights for commercial use.

## Video lecture: Produced audio: narration, podcasts and course lectures

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

1. Produced audio
2. Why produced audio differs
3. The pipeline
4. Analogy: an audiobook studio
5. Scripts for narration
6. Model choices
7. Consistency
8. QC, rights, disclosure
9. Worked example: a 17-lesson course
10. Example 2: Riyadh museum guides
11. Common mistakes
12. Watch me do it: narrate one episode
13. Recap and next step
14. Try this now

## Lecture transcript

### Produced audio

Not all voice AI is a live conversation. Course lectures, explainer videos, audiobooks, podcast segments, ads and phone menu prompts are produced audio: generated offline, reviewed and published. Latency doesn't matter here. Quality, consistency, pronunciation and rights do. In fact, the lecture you're listening to right now was produced this way, with an AI voice over generated visuals. In this lesson you'll learn the production pipeline, script writing for narration, model choices, consistency, quality control and disclosure.

### Why produced audio differs

Why treat produced audio as its own discipline? Because the audience listens longer and more carefully than a caller does. A lecture or podcast exposes every mispronunciation, every awkward pause and every change in tone across episodes. The good news is that the tools make studio-quality narration possible for small teams, if you follow a production process instead of generating everything in one go.

### The pipeline

The pipeline runs like this. Write the script for the ear. Do a pronunciation pass. Choose the voice and model. Generate in segments. Listen and fix. Assemble with music and visuals. Normalize loudness. Create captions and a transcript. Add disclosure and metadata. Publish. Generating in segments, like one scene or paragraph at a time, is the key to cheap fixes: you only regenerate the bit that went wrong.

### Analogy: an audiobook studio

An analogy: producing AI narration is like recording an audiobook in a studio, except the narrator never gets tired. The producer still prepares the script, marks pronunciations, chooses the take, listens back, and fixes individual lines rather than re-recording whole chapters. Treat your AI voice the same way: script preparation, pronunciation notes, locked settings, segment-by-segment generation, and careful listening before release.

### Scripts for narration

Scripts for narration follow a few rules. One idea per sentence, with varied sentence length for rhythm. A hook in the first fifteen seconds and a clear promise of what the listener will be able to do. Signposting, like three steps, first. Numbers, units and acronyms written the way they should be spoken, or handled with pronunciation rules. Never read URLs aloud; say the link in the lesson notes. And for teaching, concrete examples and a recap with a next step.

### Model choices

For model choice: Eleven v3 gives expressive performance, with audio tags to direct delivery, though you should test how tags behave with your voice. Multilingual v2 is known for stability across long content, which matters for a whole course or audiobook. Use dialogue features or separate voices for podcast-style segments. And for localization, dubbing tools translate and re-voice content while trying to keep speaker characteristics, but always have native speakers review the result.

### Consistency

Consistency makes a series sound professional. Fix the voice I D, model, stability, similarity and speed settings, and output format in a config file. Generate per scene, and keep a manifest mapping each script segment to its audio file. Use continuity features where available so adjacent segments sound joined up. And keep a pronunciation dictionary for brand names, technical terms and people's names. Change settings mid-series and your lessons will sound like different narrators.

### QC, rights, disclosure

Quality control is non-negotiable. Listen to everything before you publish: mispronunciations, odd emphasis, clipped endings, repeated words and artifacts. Normalize loudness to your platform's recommendation. Create captions from the script, or with speech to text, and correct them, because captions help accessibility and search. On rights: use library voices within their terms, designed voices, or consented clones. Disclose AI narration where required, and where your audience would reasonably expect to know. A line like narrated with an AI voice in the description is good practice. And check each podcast or audiobook platform's AI policy.

### Worked example: a 17-lesson course

Here's how a seventeen-lesson course gets narrated. Scripts are written per scene. A designed calm expert voice is chosen, and the model and settings are locked in a config file. A pronunciation dictionary covers terms like G A four and brand names. Scenes are generated in batches with retries, and a manifest records each scene, a hash of its text, the audio file and duration. A listening pass flags problems, and only flagged scenes are regenerated. Then the audio is assembled with visuals, normalized, captioned and published with an AI-narrated note. The Python batch script is in the lesson text.

### Example 2: Riyadh museum guides

A simple example. A Riyadh museum wants audio guides for twenty exhibits in Arabic and English. They write short scripts per exhibit, use one designed voice per language, lock the settings, and add pronunciation rules for artist and place names. A native speaker listens to every Arabic file. They generate per exhibit, so correcting one date means regenerating one file, not twenty. The guides carry a small note that narration uses AI voices.

### Common mistakes

Common production mistakes. Changing voice settings mid-series, so episodes sound like different narrators. Skipping the listening pass and publishing mispronunciations. Using a voice without clear rights for commercial use. And writing scripts for the eye, with long sentences, symbols and URLs, then wondering why the narration sounds unnatural.

### Watch me do it: narrate one episode

Watch me do it. Nova Dental wants a short patient education series, so I narrate the first episode. I split the script into six scenes, each forty to two hundred and sixty words, and save them in lecture scenes dot json with IDs. I create voice config dot json with the designed voice I D, the Multilingual v2 model for long-form stability, and the output format. I add pronunciation rules for fluoride varnish, the branch names and the dentist's surname. I run the batch script. Six MP3 files appear in the audio folder, and the manifest maps each scene I D to its file. Now the listening pass, on headphones and then on a phone speaker. Scene three rushes the word interdental, and scene five has an odd emphasis on the word only. I fix scene three with a pronunciation rule and rewrite scene five's sentence so the emphasis falls naturally. I rerun the script: only scenes three and five regenerate, because their text hashes changed, and the other four are reused untouched. Then I assemble the audio over the slides, normalize loudness to the platform's recommendation, generate captions from the script, and add a description line saying the episode is narrated with an AI voice. Total hands-on time: under an hour for a six-minute episode.

### Recap and next step

Recap. Produced audio is about quality, consistency and rights, not latency. Write for the ear, lock your voice configuration, generate in segments with a manifest, listen to everything, normalize, caption and disclose. Your next step: take one lesson you've written, split it into scenes, and narrate it with the batch script in the lesson text. Then edit one scene's text and rerun it. Only that scene should regenerate.

### Try this now

Try this now. Pick one lesson or episode and split the script into scenes of forty to two hundred and sixty words. Create a voice config file with your voice, model and settings. Create a scenes file. Run the batch script from the lesson text. Listen to every scene, flag problems, fix the text or pronunciation rules, and rerun. Only the changed scenes should regenerate. Then assemble, normalize loudness, add captions, and add an AI narration note.

## Key takeaways

- Produced audio (lectures, explainers, podcasts, ads, IVR prompts) prioritizes quality, consistency, pronunciation and rights over latency.
- Use a pipeline: script for the ear, pronunciation pass, locked voice config, per-segment generation, listening QA, loudness, captions, disclosure.
- Choose expressive models (Eleven v3) for performance and stable models (Multilingual v2) for long series; have native speakers review dubbing.
- Use voices you have rights to and disclose AI narration where required or expected.

## Try it

Split one lesson into scenes, create voice_config.json and lecture_scenes.json, and narrate it with the batch script. Edit one scene and confirm only that scene regenerates.

- [Previous: Use cases: sales qualification, support and bookings](https://optimizeall.com/learn/voice-ai-agents/use-cases-sales-support-bookings)
- [Next: Capstone: build an appointment-booking voice agent](https://optimizeall.com/learn/voice-ai-agents/capstone-appointment-booking-agent)
- [All lessons of Voice AI & Conversational Agents](https://optimizeall.com/learn/voice-ai-agents)
