AI Video & Voice Production: ElevenLabs, Veo, Runway and MoreProduction playbooks: course videos and marketing videos · Lesson 15 of 16
Playbook: producing course and training videos with AI narration
Video lecture
Playbook: producing course and training videos with AI narration
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 The course-video playbook
Here's a confession about the lecture you're watching right now. There's no presenter, no camera, and no avatar. There's a script written as scenes, a narrator voice generated per scene, and animated slides timed to the words. That format is one of the most scalable ways to teach in twenty twenty-six. In this lecture I'll show you the full pipeline behind it: how to write scenes that are actually producible, how to voice them automatically, how editors sync visuals to narration, and how to keep a course accurate after launch.
0:39 Why this format
Why this format? Because teaching has different needs from marketing. Learners need to see diagrams, code and data, not a face. Courses need frequent updates when tools rename features or regulations change. And many courses need several languages. Narration over slides handles all three. You can fix one scene without re-shooting, localize by swapping voice and text, and keep the visual channel entirely for what learners need to understand. And there's no uncanny avatar to distract anyone.
1:12 The comic-book model
Here's the mental model. Think of each lecture as a comic book with a soundtrack. Each panel is a scene. The speech bubble is the narration, the caption box is the on-screen text, and the artist's notes say what's drawn and how it moves. If the words, the caption and the drawing don't match, the panel confuses the reader. In our case, every scene needs three matching parts: what's said, what's shown as text, and what moves on screen, and when.
1:47 The eight-step pipeline
Here's the pipeline in eight steps. One, a learning objective in one sentence. Two, the lecture script as scenes: narration, on-screen text, visual direction and seconds. Three, expert review of facts, code and claims. That's gate one. Four, voice: a consistent narrator, per-scene audio, and a pronunciation list. Five, slides and visuals from a template. Six, assembly, where the audio drives timing. Seven, captions and quality control, gate two. And eight, publish and maintain, with a version, a changelog and a review date.
2:23 Producible visual direction
Let's zoom in on the visual direction, because it's what makes a scene producible. It should read like instructions to a motion designer. Start with the layout: two tables on the left, a result table on the right. Then timing beats. Zero to fifteen seconds, the orders table draws. Fifteen to thirty, the items table draws and three connector lines fan out. Thirty to forty-five, the result fills and the total turns red. Then a callout slides in. With that, an editor can sync animation to narration without guessing.
3:02 The teaching arc
And every lecture follows a teaching arc. A hook and a promise of what you'll be able to do. Why it matters. The concept, with an analogy. A simple worked example. A realistic business example. A watch me do it walkthrough. Common mistakes. A recap. And a try this now action. You've heard that arc in every lecture of this course. It works because it moves from curiosity to understanding to practice, and learners always leave with a next step.
3:37 Pace and load
A few rules for pace and load. Narration at roughly one hundred thirty to one hundred fifty words per minute feels calm for teaching, so an eight to ten minute lecture is about eleven hundred to fourteen hundred words. One idea per scene. On-screen text supports the narration and never duplicates it word for word, with four short bullets at most. And show code only as large, highlighted snippets synced with what's being said, with the full code in the lesson text.
4:13 Example 1: monthly academy releases
First example, a simple one. An online academy publishes twelve lectures a month. Each course category has a fixed narrator voice, a slide template with light and dark variants, and a pronunciation list per course. When a platform renames a feature, the team edits one scene's narration and on-screen text, regenerates one audio file, re-exports, and updates the last reviewed date on the course page. A change that used to mean a re-shoot now takes under an hour.
4:47 Example 2: Riyadh bank academy
Second example, a business case. A bank's learning team in Riyadh converts a forty-page compliance manual into eight short lectures in Arabic and English. Each lecture opens with a real, anonymized scenario from the bank's call center. Procedures become animated flowcharts. Every lecture ends with a try this now task in the internal system. And they use a designed Arabic narrator and a designed English narrator, so the course sounds consistent without tying it to any one employee who might leave.
5:22 Watch me do it, part 1
Watch me do it. I keep each lecture's scenes in a JSON file, which is also how this academy stores them. The lesson's script loads that file, and for each scene it checks whether the audio already exists. If not, it calls text to speech with my narrator voice and Multilingual v2 for stable, long-form delivery, and saves a file named by lecture and scene number. If a scene fails, it prints the error and moves on, so one problem doesn't stop the batch.
5:59 Watch me do it, part 2
Then it writes a timing sheet. For each scene: the audio file, the planned seconds, the on-screen text and the visual direction. The editor, or a slide-animation tool, reads that sheet and builds each scene. The real audio length replaces the planned seconds, and the animation beats are stretched or squeezed to match. Finally, I transcribe the assembled lecture for captions, run the QC checklist, and add an entry to the lecture's changelog with today's date.
6:32 Maintenance
Now maintenance, the part most teams forget. Courses decay. Tools rename features, prices change, laws take effect. So store scripts as data, not only as videos. Keep a changelog per lecture. Set a review date and show it to learners. And watch learner questions. If people keep asking about the same scene, that scene is confusing, and you can fix it in minutes because it's just text and one audio file.
7:03 Common mistakes
Common mistakes. Reading the lesson text aloud instead of teaching it in spoken form with its own examples. Learners can read; the lecture has to teach. Dense slides that compete with narration. No timing notes, so editors guess where animations go. Voicing before expert review, then re-voicing after corrections. And no review date, so learners can't tell whether content is current.
7:30 Recap
Recap. Narration over animated slides scales beautifully for teaching. Every scene needs matching narration, on-screen text and a visual direction with timing. Follow the teaching arc. Put expert review before voicing, QC before publishing, and keep scripts as data so updates take minutes.
7:49 Try this now
Try this now. Pick one topic you can teach in eight minutes. Write the one-sentence objective, then draft ten scenes with narration, on-screen text and a visual direction that starts with the layout and includes timing beats. Read two scenes aloud with a timer, and ask a colleague whether they could animate the visual directions without asking you a single question.
The course-video format that scales
For teaching, the most scalable format in 2026 is AI narration over animated slides: an expert-reviewed script voiced with a consistent narrator, over branded slides with diagrams, code walkthroughs, charts and key-idea cards. No presenter on camera, no avatar. It is quick to update when facts change, easy to localize, and it keeps the visual channel free for what learners need to see. This academy's own lectures follow this pattern.
The course-video pipeline
1. Learning objective -> one sentence: "After this lecture you can ..."
2. Lecture script -> scenes: narration | on-screen text | visual direction | seconds
3. Expert review -> facts, code and claims checked (GATE 1)
4. Voice -> designed or licensed narrator, per-scene audio, pronunciation list
5. Slides/visuals -> template + per-scene layouts, code and diagrams, generated b-roll only for mood
6. Assembly -> audio drives timing; visuals follow the "visual" timing notes
7. Captions + QC -> transcript from final audio; checklist; phone test (GATE 2)
8. Publish + maintain -> version, changelog, review date ("last reviewed: 2026-09")Writing lecture scenes that are producible
Every scene needs three things that match: what is said, what is shown as text, and what moves on screen, when. A good visual direction reads like instructions to a motion designer:
Scene 6 (52 s)
Narration: "... the join multiplies rows, so the sum is inflated ..."
On-screen: "Join fan-out\n• One order, three items\n• Total counted three times"
Visual: Layout: two tables left, result table right. 0-15 s orders table draws;
15-30 s items table draws and three connector lines fan out from order 1001;
30-45 s result table fills with 1001 repeated three times, total cell turns red;
45-52 s callout "aggregate first" slides in.Timing notes let the editor sync animation to narration without guessing. Aim for a teaching arc in every lecture: hook, why it matters, concept with an analogy, a simple worked example, a realistic business example, a "watch me do it" walkthrough, common mistakes, recap and a "try this now" action.
Pace, length and cognitive load
- Narration at roughly 130 to 150 words per minute feels calm for teaching; an 8 to 10 minute lecture is about 1,100 to 1,400 words.
- One idea per scene, 40 to 150 seconds each.
- On-screen text supports narration; it does not duplicate it. Four short bullets at most.
- Show code only as large, highlighted snippets synchronized with narration; link to the full code in the lesson text.
Hands-on: generate a lecture from a scene file
Keep lecture scenes in a JSON or CSV file (this is also how this academy stores them). A short script can voice every scene and produce a timing sheet for the editor:
import json, os
from elevenlabs.client import ElevenLabs
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
VOICE_ID, MODEL_ID = os.environ["NARRATOR_VOICE_ID"], "eleven_multilingual_v2"
lecture = json.load(open("lecture_07.json", encoding="utf-8"))
timing = []
for i, scene in enumerate(lecture["scenes"], start=1):
path = f"l07_s{i:02d}.mp3"
if not os.path.exists(path):
try:
audio = client.text_to_speech.convert(text=scene["narration"], voice_id=VOICE_ID,
model_id=MODEL_ID, output_format="mp3_44100_128")
with open(path, "wb") as f:
for chunk in audio:
f.write(chunk)
except Exception as err:
print("scene", i, "failed:", err)
continue
timing.append({"scene": i, "audio": path, "planned_seconds": scene["seconds"],
"on_screen": scene["onScreen"], "visual": scene["visual"]})
json.dump(timing, open("l07_timing.json", "w", encoding="utf-8"), indent=2, ensure_ascii=False)The editor (or a slide-animation tool) uses l07_timing.json to build each scene; the real audio length replaces the planned seconds.
Worked example: an academy's monthly release
An online academy publishes twelve lectures a month across AI, marketing and business courses. Each course category has a fixed narrator voice (for example a US male voice for AI and data, a US female voice for marketing, a British male voice for sales), a slide template with light and dark variants, and a pronunciation list per course. Experts review scripts before voicing. When a platform renames a feature, the team edits one scene's narration and on-screen text, regenerates one audio file, and re-exports; the "last reviewed" date is updated on the course page.
Second worked example: a corporate academy in Riyadh
A bank's learning team converts a 40-page compliance manual into eight short lectures in Arabic and English. Each lecture opens with a real scenario from the bank's call center (anonymized), uses animated flowcharts for procedures, and ends with a "try this now" task in the internal system. A designed Arabic narrator and a designed English narrator keep consistency without tying the course to one employee.
Maintenance: the part most teams forget
Courses decay. Store scripts as data, keep a changelog per lecture, set a review date, and track learner questions: repeated confusion at one scene is a signal to rewrite it.
Pitfalls
- Reading the lesson text aloud instead of teaching it in spoken form with its own examples.
- Dense slides that compete with narration.
- No timing notes, so editors guess where animations go.
- Voicing before expert review, then re-voicing after corrections.
Key takeaways
- AI narration over animated slides is a highly scalable teaching format: easy to update, localize and keep accurate.
- Every scene needs matching narration, on-screen text and a visual direction that starts with the layout and includes timing.
- Follow a teaching arc: hook, why, concept, two examples, walkthrough, mistakes, recap, try this now.
- Store scripts as data with changelogs and review dates, so updates take minutes, not re-shoots.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Write a one-sentence objective and ten producible scenes for an eight-minute lecture, then read two scenes aloud with a timer.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.