---
title: "Playbook: producing course and training videos with AI…"
description: "The course-video format that scales For teaching, the most scalable format in 2026 is AI narration over animated slides : an expert-reviewed script…"
url: https://optimizeall.com/learn/ai-video-and-voice-production/course-video-production-playbook
updated: 2026-10-05
---

AI Video & Voice Production: ElevenLabs, Veo, Runway and More · Production playbooks: course videos and marketing videos · lesson 15 of 16 · 8 min

# Playbook: producing course and training videos with AI narration

## The course-video format that scales

For teaching, the most scalable format in 2026 is **AI narration over animated slides**: an expert-reviewed script voiced with a consistent narrator, over branded slides with diagrams, code walkthroughs, charts and key-idea cards. No presenter on camera, no avatar. It is quick to update when facts change, easy to localize, and it keeps the visual channel free for what learners need to see. This academy's own lectures follow this pattern.

## The course-video pipeline

```text
1. Learning objective   -> one sentence: "After this lecture you can ..."
2. Lecture script        -> scenes: narration | on-screen text | visual direction | seconds
3. Expert review         -> facts, code and claims checked (GATE 1)
4. Voice                 -> designed or licensed narrator, per-scene audio, pronunciation list
5. Slides/visuals        -> template + per-scene layouts, code and diagrams, generated b-roll only for mood
6. Assembly              -> audio drives timing; visuals follow the "visual" timing notes
7. Captions + QC         -> transcript from final audio; checklist; phone test (GATE 2)
8. Publish + maintain    -> version, changelog, review date ("last reviewed: 2026-09")
```

## Writing lecture scenes that are producible

Every scene needs three things that match: what is **said**, what is **shown as text**, and what **moves on screen, when**. A good visual direction reads like instructions to a motion designer:

```text
Scene 6 (52 s)
Narration:  "... the join multiplies rows, so the sum is inflated ..."
On-screen:  "Join fan-out\n• One order, three items\n• Total counted three times"
Visual:     Layout: two tables left, result table right. 0-15 s orders table draws;
            15-30 s items table draws and three connector lines fan out from order 1001;
            30-45 s result table fills with 1001 repeated three times, total cell turns red;
            45-52 s callout "aggregate first" slides in.
```

Timing notes let the editor sync animation to narration without guessing. Aim for a teaching arc in every lecture: hook, why it matters, concept with an analogy, a simple worked example, a realistic business example, a "watch me do it" walkthrough, common mistakes, recap and a "try this now" action.

## Pace, length and cognitive load

- Narration at roughly 130 to 150 words per minute feels calm for teaching; an 8 to 10 minute lecture is about 1,100 to 1,400 words.
- One idea per scene, 40 to 150 seconds each.
- On-screen text supports narration; it does not duplicate it. Four short bullets at most.
- Show code only as large, highlighted snippets synchronized with narration; link to the full code in the lesson text.

## Hands-on: generate a lecture from a scene file

Keep lecture scenes in a JSON or CSV file (this is also how this academy stores them). A short script can voice every scene and produce a timing sheet for the editor:

```python
import json, os
from elevenlabs.client import ElevenLabs

client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
VOICE_ID, MODEL_ID = os.environ["NARRATOR_VOICE_ID"], "eleven_multilingual_v2"

lecture = json.load(open("lecture_07.json", encoding="utf-8"))
timing = []
for i, scene in enumerate(lecture["scenes"], start=1):
    path = f"l07_s{i:02d}.mp3"
    if not os.path.exists(path):
        try:
            audio = client.text_to_speech.convert(text=scene["narration"], voice_id=VOICE_ID,
                                                  model_id=MODEL_ID, output_format="mp3_44100_128")
            with open(path, "wb") as f:
                for chunk in audio:
                    f.write(chunk)
        except Exception as err:
            print("scene", i, "failed:", err)
            continue
    timing.append({"scene": i, "audio": path, "planned_seconds": scene["seconds"],
                   "on_screen": scene["onScreen"], "visual": scene["visual"]})

json.dump(timing, open("l07_timing.json", "w", encoding="utf-8"), indent=2, ensure_ascii=False)
```

The editor (or a slide-animation tool) uses `l07_timing.json` to build each scene; the real audio length replaces the planned seconds.

## Worked example: an academy's monthly release

An online academy publishes twelve lectures a month across AI, marketing and business courses. Each course category has a fixed narrator voice (for example a US male voice for AI and data, a US female voice for marketing, a British male voice for sales), a slide template with light and dark variants, and a pronunciation list per course. Experts review scripts before voicing. When a platform renames a feature, the team edits one scene's narration and on-screen text, regenerates one audio file, and re-exports; the "last reviewed" date is updated on the course page.

## Second worked example: a corporate academy in Riyadh

A bank's learning team converts a 40-page compliance manual into eight short lectures in Arabic and English. Each lecture opens with a real scenario from the bank's call center (anonymized), uses animated flowcharts for procedures, and ends with a "try this now" task in the internal system. A designed Arabic narrator and a designed English narrator keep consistency without tying the course to one employee.

## Maintenance: the part most teams forget

Courses decay. Store scripts as data, keep a changelog per lecture, set a review date, and track learner questions: repeated confusion at one scene is a signal to rewrite it.

## Pitfalls

- Reading the lesson text aloud instead of teaching it in spoken form with its own examples.
- Dense slides that compete with narration.
- No timing notes, so editors guess where animations go.
- Voicing before expert review, then re-voicing after corrections.

## Video lecture: Playbook: producing course and training videos with AI narration

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. The course-video playbook
2. Why this format
3. The comic-book model
4. The eight-step pipeline
5. Producible visual direction
6. The teaching arc
7. Pace and load
8. Example 1: monthly academy releases
9. Example 2: Riyadh bank academy
10. Watch me do it, part 1
11. Watch me do it, part 2
12. Maintenance
13. Common mistakes
14. Recap
15. Try this now

## Lecture transcript

### The course-video playbook

Here's a confession about the lecture you're watching right now. There's no presenter, no camera, and no avatar. There's a script written as scenes, a narrator voice generated per scene, and animated slides timed to the words. That format is one of the most scalable ways to teach in twenty twenty-six. In this lecture I'll show you the full pipeline behind it: how to write scenes that are actually producible, how to voice them automatically, how editors sync visuals to narration, and how to keep a course accurate after launch.

### Why this format

Why this format? Because teaching has different needs from marketing. Learners need to see diagrams, code and data, not a face. Courses need frequent updates when tools rename features or regulations change. And many courses need several languages. Narration over slides handles all three. You can fix one scene without re-shooting, localize by swapping voice and text, and keep the visual channel entirely for what learners need to understand. And there's no uncanny avatar to distract anyone.

### The comic-book model

Here's the mental model. Think of each lecture as a comic book with a soundtrack. Each panel is a scene. The speech bubble is the narration, the caption box is the on-screen text, and the artist's notes say what's drawn and how it moves. If the words, the caption and the drawing don't match, the panel confuses the reader. In our case, every scene needs three matching parts: what's said, what's shown as text, and what moves on screen, and when.

### The eight-step pipeline

Here's the pipeline in eight steps. One, a learning objective in one sentence. Two, the lecture script as scenes: narration, on-screen text, visual direction and seconds. Three, expert review of facts, code and claims. That's gate one. Four, voice: a consistent narrator, per-scene audio, and a pronunciation list. Five, slides and visuals from a template. Six, assembly, where the audio drives timing. Seven, captions and quality control, gate two. And eight, publish and maintain, with a version, a changelog and a review date.

### Producible visual direction

Let's zoom in on the visual direction, because it's what makes a scene producible. It should read like instructions to a motion designer. Start with the layout: two tables on the left, a result table on the right. Then timing beats. Zero to fifteen seconds, the orders table draws. Fifteen to thirty, the items table draws and three connector lines fan out. Thirty to forty-five, the result fills and the total turns red. Then a callout slides in. With that, an editor can sync animation to narration without guessing.

### The teaching arc

And every lecture follows a teaching arc. A hook and a promise of what you'll be able to do. Why it matters. The concept, with an analogy. A simple worked example. A realistic business example. A watch me do it walkthrough. Common mistakes. A recap. And a try this now action. You've heard that arc in every lecture of this course. It works because it moves from curiosity to understanding to practice, and learners always leave with a next step.

### Pace and load

A few rules for pace and load. Narration at roughly one hundred thirty to one hundred fifty words per minute feels calm for teaching, so an eight to ten minute lecture is about eleven hundred to fourteen hundred words. One idea per scene. On-screen text supports the narration and never duplicates it word for word, with four short bullets at most. And show code only as large, highlighted snippets synced with what's being said, with the full code in the lesson text.

### Example 1: monthly academy releases

First example, a simple one. An online academy publishes twelve lectures a month. Each course category has a fixed narrator voice, a slide template with light and dark variants, and a pronunciation list per course. When a platform renames a feature, the team edits one scene's narration and on-screen text, regenerates one audio file, re-exports, and updates the last reviewed date on the course page. A change that used to mean a re-shoot now takes under an hour.

### Example 2: Riyadh bank academy

Second example, a business case. A bank's learning team in Riyadh converts a forty-page compliance manual into eight short lectures in Arabic and English. Each lecture opens with a real, anonymized scenario from the bank's call center. Procedures become animated flowcharts. Every lecture ends with a try this now task in the internal system. And they use a designed Arabic narrator and a designed English narrator, so the course sounds consistent without tying it to any one employee who might leave.

### Watch me do it, part 1

Watch me do it. I keep each lecture's scenes in a JSON file, which is also how this academy stores them. The lesson's script loads that file, and for each scene it checks whether the audio already exists. If not, it calls text to speech with my narrator voice and Multilingual v2 for stable, long-form delivery, and saves a file named by lecture and scene number. If a scene fails, it prints the error and moves on, so one problem doesn't stop the batch.

### Watch me do it, part 2

Then it writes a timing sheet. For each scene: the audio file, the planned seconds, the on-screen text and the visual direction. The editor, or a slide-animation tool, reads that sheet and builds each scene. The real audio length replaces the planned seconds, and the animation beats are stretched or squeezed to match. Finally, I transcribe the assembled lecture for captions, run the QC checklist, and add an entry to the lecture's changelog with today's date.

### Maintenance

Now maintenance, the part most teams forget. Courses decay. Tools rename features, prices change, laws take effect. So store scripts as data, not only as videos. Keep a changelog per lecture. Set a review date and show it to learners. And watch learner questions. If people keep asking about the same scene, that scene is confusing, and you can fix it in minutes because it's just text and one audio file.

### Common mistakes

Common mistakes. Reading the lesson text aloud instead of teaching it in spoken form with its own examples. Learners can read; the lecture has to teach. Dense slides that compete with narration. No timing notes, so editors guess where animations go. Voicing before expert review, then re-voicing after corrections. And no review date, so learners can't tell whether content is current.

### Recap

Recap. Narration over animated slides scales beautifully for teaching. Every scene needs matching narration, on-screen text and a visual direction with timing. Follow the teaching arc. Put expert review before voicing, QC before publishing, and keep scripts as data so updates take minutes.

### Try this now

Try this now. Pick one topic you can teach in eight minutes. Write the one-sentence objective, then draft ten scenes with narration, on-screen text and a visual direction that starts with the layout and includes timing beats. Read two scenes aloud with a timer, and ask a colleague whether they could animate the visual directions without asking you a single question.

## Key takeaways

- AI narration over animated slides is a highly scalable teaching format: easy to update, localize and keep accurate.
- Every scene needs matching narration, on-screen text and a visual direction that starts with the layout and includes timing.
- Follow a teaching arc: hook, why, concept, two examples, walkthrough, mistakes, recap, try this now.
- Store scripts as data with changelogs and review dates, so updates take minutes, not re-shoots.

## Try it

Write a one-sentence objective and ten producible scenes for an eight-minute lecture, then read two scenes aloud with a timer.

- [Previous: Disclosing synthetic media and following platform policies](https://optimizeall.com/learn/ai-video-and-voice-production/disclosure-and-platform-policies)
- [Next: Playbook: marketing videos, variants and honest measurement](https://optimizeall.com/learn/ai-video-and-voice-production/marketing-video-playbook)
- [All lessons of AI Video & Voice Production: ElevenLabs, Veo, Runway and More](https://optimizeall.com/learn/ai-video-and-voice-production)
