AI Video & Voice Production: ElevenLabs, Veo, Runway and MoreProduction playbooks: course videos and marketing videos · Lesson 15 of 16

Playbook: producing course and training videos with AI narration

Article · 8 min · 8 min lecture

Video lecture

Playbook: producing course and training videos with AI narration

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

The course-video playbook

  • Narration over animated slides
  • Producible lecture scenes
  • Automated voicing and timing
  • Keeping courses current

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The course-video format that scales

For teaching, the most scalable format in 2026 is AI narration over animated slides: an expert-reviewed script voiced with a consistent narrator, over branded slides with diagrams, code walkthroughs, charts and key-idea cards. No presenter on camera, no avatar. It is quick to update when facts change, easy to localize, and it keeps the visual channel free for what learners need to see. This academy's own lectures follow this pattern.

The course-video pipeline

1. Learning objective   -> one sentence: "After this lecture you can ..."
2. Lecture script        -> scenes: narration | on-screen text | visual direction | seconds
3. Expert review         -> facts, code and claims checked (GATE 1)
4. Voice                 -> designed or licensed narrator, per-scene audio, pronunciation list
5. Slides/visuals        -> template + per-scene layouts, code and diagrams, generated b-roll only for mood
6. Assembly              -> audio drives timing; visuals follow the "visual" timing notes
7. Captions + QC         -> transcript from final audio; checklist; phone test (GATE 2)
8. Publish + maintain    -> version, changelog, review date ("last reviewed: 2026-09")

Writing lecture scenes that are producible

Every scene needs three things that match: what is said, what is shown as text, and what moves on screen, when. A good visual direction reads like instructions to a motion designer:

Scene 6 (52 s)
Narration:  "... the join multiplies rows, so the sum is inflated ..."
On-screen:  "Join fan-out\n• One order, three items\n• Total counted three times"
Visual:     Layout: two tables left, result table right. 0-15 s orders table draws;
            15-30 s items table draws and three connector lines fan out from order 1001;
            30-45 s result table fills with 1001 repeated three times, total cell turns red;
            45-52 s callout "aggregate first" slides in.

Timing notes let the editor sync animation to narration without guessing. Aim for a teaching arc in every lecture: hook, why it matters, concept with an analogy, a simple worked example, a realistic business example, a "watch me do it" walkthrough, common mistakes, recap and a "try this now" action.

Pace, length and cognitive load

  • Narration at roughly 130 to 150 words per minute feels calm for teaching; an 8 to 10 minute lecture is about 1,100 to 1,400 words.
  • One idea per scene, 40 to 150 seconds each.
  • On-screen text supports narration; it does not duplicate it. Four short bullets at most.
  • Show code only as large, highlighted snippets synchronized with narration; link to the full code in the lesson text.

Hands-on: generate a lecture from a scene file

Keep lecture scenes in a JSON or CSV file (this is also how this academy stores them). A short script can voice every scene and produce a timing sheet for the editor:

import json, os
from elevenlabs.client import ElevenLabs

client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
VOICE_ID, MODEL_ID = os.environ["NARRATOR_VOICE_ID"], "eleven_multilingual_v2"

lecture = json.load(open("lecture_07.json", encoding="utf-8"))
timing = []
for i, scene in enumerate(lecture["scenes"], start=1):
    path = f"l07_s{i:02d}.mp3"
    if not os.path.exists(path):
        try:
            audio = client.text_to_speech.convert(text=scene["narration"], voice_id=VOICE_ID,
                                                  model_id=MODEL_ID, output_format="mp3_44100_128")
            with open(path, "wb") as f:
                for chunk in audio:
                    f.write(chunk)
        except Exception as err:
            print("scene", i, "failed:", err)
            continue
    timing.append({"scene": i, "audio": path, "planned_seconds": scene["seconds"],
                   "on_screen": scene["onScreen"], "visual": scene["visual"]})

json.dump(timing, open("l07_timing.json", "w", encoding="utf-8"), indent=2, ensure_ascii=False)

The editor (or a slide-animation tool) uses l07_timing.json to build each scene; the real audio length replaces the planned seconds.

Worked example: an academy's monthly release

An online academy publishes twelve lectures a month across AI, marketing and business courses. Each course category has a fixed narrator voice (for example a US male voice for AI and data, a US female voice for marketing, a British male voice for sales), a slide template with light and dark variants, and a pronunciation list per course. Experts review scripts before voicing. When a platform renames a feature, the team edits one scene's narration and on-screen text, regenerates one audio file, and re-exports; the "last reviewed" date is updated on the course page.

Second worked example: a corporate academy in Riyadh

A bank's learning team converts a 40-page compliance manual into eight short lectures in Arabic and English. Each lecture opens with a real scenario from the bank's call center (anonymized), uses animated flowcharts for procedures, and ends with a "try this now" task in the internal system. A designed Arabic narrator and a designed English narrator keep consistency without tying the course to one employee.

Maintenance: the part most teams forget

Courses decay. Store scripts as data, keep a changelog per lecture, set a review date, and track learner questions: repeated confusion at one scene is a signal to rewrite it.

Pitfalls

  • Reading the lesson text aloud instead of teaching it in spoken form with its own examples.
  • Dense slides that compete with narration.
  • No timing notes, so editors guess where animations go.
  • Voicing before expert review, then re-voicing after corrections.

Key takeaways

  • AI narration over animated slides is a highly scalable teaching format: easy to update, localize and keep accurate.
  • Every scene needs matching narration, on-screen text and a visual direction that starts with the layout and includes timing.
  • Follow a teaching arc: hook, why, concept, two examples, walkthrough, mistakes, recap, try this now.
  • Store scripts as data with changelogs and review dates, so updates take minutes, not re-shoots.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which visual direction is most producible for a motion designer?
  2. A tool renamed a feature mentioned in lecture 7. What does the scripts-as-data approach allow?

Put it into practice

Write a one-sentence objective and ten producible scenes for an eight-minute lecture, then read two scenes aloud with a timer.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.