AI Video & Voice Production: ElevenLabs, Veo, Runway and MoreScripting and production workflows · Lesson 8 of 16

Scripting for AI voices and avatars

Article · 10 min · 8 min lecture

Video lecture

Scripting for AI voices and avatars

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

Scripting for AI voice and avatars

  • Writing for the ear
  • Sizing scripts to time
  • The scene-based format
  • AI-drafted scripts you can produce

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Writing for the ear, not the eye

Scripts for AI voices and avatars must be written to be heard. A sentence that reads well on a page can sound robotic when spoken, especially by a synthetic voice that performs exactly what is written.

Core rules

  1. Short sentences. Aim for one idea per sentence. Long, clause-heavy sentences flatten intonation.
  2. Conversational words. "Use" not "utilize". "Help" not "facilitate". Contractions sound natural.
  3. Front-load the point. Viewers decide in the first seconds. Open with the hook, not a greeting.
  4. Write numbers the way you want them spoken. "Two thousand five hundred rupees" versus "Rs 2,500". Test which your tool reads correctly.
  5. Spell out acronyms and brand names as they should sound: "S-E-O" or "seo", "HeyGen" as "Hey Jen".
  6. Use punctuation as direction. Commas for breath, full stops for beats, a new paragraph for a longer pause. Some tools support explicit pause markup; check your tool.
  7. Avoid tongue-twisters and dense lists. Three items read aloud is plenty.
  8. Mark emphasis sparingly. Rephrasing ("Here's the key part:") is often more reliable than formatting tricks.

Pacing guide

Comfortable narration typically runs at about 130 to 160 words per minute, depending on language, voice and style. Use this to size scripts:

Video lengthApproximate words
15 seconds35–40
30 seconds65–80
60 seconds130–160
2 minutes260–320

Leave room for pauses, visuals and on-screen text.

Structure for short-form

  • Hook (0–3s): a problem, surprising statement or promise the video keeps.
  • Setup (3–10s): why it matters to this viewer.
  • Value (10–45s): two or three concrete points, each with an example.
  • CTA (last 5s): one clear action.

Scripts for avatars

Avatars add visual considerations:

  • Scene breaks: split scripts into scenes of about 10 to 20 seconds. Cut to b-roll, screen recordings or text between scenes to keep energy and hide avatar limitations.
  • Gestures and expressions: some platforms allow gesture or expression cues. Use them lightly; overuse looks artificial.
  • Eye contact and framing: keep the avatar talking to camera; use b-roll for demonstrations.
  • On-screen text: plan key phrases as captions or titles; do not rely on the avatar alone.

Two-column script format

Use this format so voice, avatar and editor stay aligned:

SceneVoiceover / avatar lineVisual
1 (0–4s)"Your Reels get views, but no DMs? Here's why."Avatar, close-up. Text: "Views ≠ DMs"
2 (4–15s)"Most captions end with 'link in bio'. That's a dead end for DM-first buyers."B-roll: phone scrolling; screenshot of generic caption (anonymized)
3 (15–35s)"Instead, give one keyword to comment. Then reply with a helpful DM, not a sales pitch."Screen recording of the comment-to-DM flow
4 (35–40s)"Comment 'DM' and I'll send you my three best CTA templates."Avatar. Text: "Comment DM"

Using AI to draft scripts

Combine this lesson with your prompt skills:

Write a 40-second vertical video script for [audience] about [topic]. Written for the ear: short sentences, contractions, no acronyms without phonetic spelling. Target 90–100 words. Format as a table with Scene, Voiceover, Visual. Split into 4 scenes. Hook in under 8 words. Do not add claims beyond these facts: [facts].

Then read it aloud yourself before generating. If you stumble, the AI voice will sound off too.

Worked example: fixing a script line

Before: "Utilizing our cutting-edge, AI-driven, multi-channel CRM (customer relationship management) platform, SMEs can streamline their omnichannel engagement workflows."

After: "Here's the problem. Your customers message you on WhatsApp, Instagram and email. And you're losing track. Our CRM puts every message in one inbox."

Shorter, spoken, concrete, and much easier for an AI voice to deliver naturally.

Hands-on: direction cues for expressive models

Expressive models such as ElevenLabs Eleven v3 accept inline audio tags in square brackets. Use them like a director's margin notes, sparingly:

[warm] Most Reels get views. [short pause] Very few get DMs.
Here's why. [conversational] Your caption ends with "link in bio" - and that's a dead end.
[emphatic] Give people one word to comment instead.

Rules of thumb: one cue per sentence at most, test every cue with your chosen voice (a cue that suits one voice can sound odd on another), and keep a clean version without tags for models that read brackets literally. For stable long-form narration, rewriting the sentence is usually more reliable than a tag.

A script prompt that produces producible scenes

You are a scriptwriter for AI-voiced explainer videos.
Audience: {who}. Platform: {platform, aspect ratio}. Length: {seconds} s (~{words} words at 150 wpm).
Goal: {one action the viewer should take}.
Facts you may use (do not add others): {bullet facts}.
Write a table with columns: Scene | Seconds | Voiceover (spoken, short sentences) | On-screen text (max 6 words) | Visual (what the editor shows).
Rules: hook in under 8 words; one idea per scene; spell acronyms phonetically; numbers written as spoken;
no claims beyond the facts; end with one CTA. Then list 3 words the voice may mispronounce.

The last line is a small trick: asking the model to predict mispronunciations gives you a head start on the pronunciation list.

Second worked example: a KSA real-estate developer

A Riyadh developer needs a 45-second Arabic and English teaser for a new community. The first AI draft is full of superlatives ("the most luxurious", "unmatched") that the legal team will not approve and that are hard to voice naturally. The rewrite uses concrete facts: distance to the metro, number of parks, handover quarter. Each scene has one fact and one visual. The Arabic version is written separately by a native writer from the same fact list rather than translated line by line, because sentence rhythm differs between languages. Both scripts are read aloud before any generation.

Pitfalls

  • Pasting blog copy straight into a voice tool.
  • Scripts too long for the target length, which forces rushed speed settings.
  • Forgetting to test pronunciation of local names such as place names, dishes and people.

Key takeaways

  • Write for the ear: short sentences, conversational words, a front-loaded hook.
  • Size scripts at roughly 130–160 spoken words per minute, leaving room for pauses.
  • Split avatar scripts into short scenes with b-roll and on-screen text between them.
  • Use a two-column Scene / Voiceover / Visual format and read scripts aloud before generating.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. About how many words suit a 30-second AI voiceover?
  2. Why split an avatar script into short scenes?

Put it into practice

Rewrite a paragraph from your website or blog as a 30-second spoken script in the two-column format, read it aloud, and fix every line you stumble on.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.