---
title: "Scripting for AI voices and avatars | Optimize All Academy"
description: "Writing for the ear, not the eye Scripts for AI voices and avatars must be written to be heard . A sentence that reads well on a page can sound robotic…"
url: https://optimizeall.com/learn/ai-video-and-voice-production/scripting-for-ai-voice-and-avatars
updated: 2026-10-05
---

AI Video & Voice Production: ElevenLabs, Veo, Runway and More · Scripting and production workflows · lesson 8 of 16 · 10 min

# Scripting for AI voices and avatars

## Writing for the ear, not the eye

Scripts for AI voices and avatars must be written to be **heard**. A sentence that reads well on a page can sound robotic when spoken, especially by a synthetic voice that performs exactly what is written.

## Core rules

1. **Short sentences.** Aim for one idea per sentence. Long, clause-heavy sentences flatten intonation.
2. **Conversational words.** "Use" not "utilize". "Help" not "facilitate". Contractions sound natural.
3. **Front-load the point.** Viewers decide in the first seconds. Open with the hook, not a greeting.
4. **Write numbers the way you want them spoken.** "Two thousand five hundred rupees" versus "Rs 2,500". Test which your tool reads correctly.
5. **Spell out acronyms and brand names** as they should sound: "S-E-O" or "seo", "HeyGen" as "Hey Jen".
6. **Use punctuation as direction.** Commas for breath, full stops for beats, a new paragraph for a longer pause. Some tools support explicit pause markup; check your tool.
7. **Avoid tongue-twisters and dense lists.** Three items read aloud is plenty.
8. **Mark emphasis sparingly.** Rephrasing ("Here's the key part:") is often more reliable than formatting tricks.

## Pacing guide

Comfortable narration typically runs at about 130 to 160 words per minute, depending on language, voice and style. Use this to size scripts:

| Video length | Approximate words |
|---|---|
| 15 seconds | 35–40 |
| 30 seconds | 65–80 |
| 60 seconds | 130–160 |
| 2 minutes | 260–320 |

Leave room for pauses, visuals and on-screen text.

## Structure for short-form

- **Hook (0–3s):** a problem, surprising statement or promise the video keeps.
- **Setup (3–10s):** why it matters to this viewer.
- **Value (10–45s):** two or three concrete points, each with an example.
- **CTA (last 5s):** one clear action.

## Scripts for avatars

Avatars add visual considerations:

- **Scene breaks:** split scripts into scenes of about 10 to 20 seconds. Cut to b-roll, screen recordings or text between scenes to keep energy and hide avatar limitations.
- **Gestures and expressions:** some platforms allow gesture or expression cues. Use them lightly; overuse looks artificial.
- **Eye contact and framing:** keep the avatar talking to camera; use b-roll for demonstrations.
- **On-screen text:** plan key phrases as captions or titles; do not rely on the avatar alone.

## Two-column script format

Use this format so voice, avatar and editor stay aligned:

| Scene | Voiceover / avatar line | Visual |
|---|---|---|
| 1 (0–4s) | "Your Reels get views, but no DMs? Here's why." | Avatar, close-up. Text: "Views ≠ DMs" |
| 2 (4–15s) | "Most captions end with 'link in bio'. That's a dead end for DM-first buyers." | B-roll: phone scrolling; screenshot of generic caption (anonymized) |
| 3 (15–35s) | "Instead, give one keyword to comment. Then reply with a helpful DM, not a sales pitch." | Screen recording of the comment-to-DM flow |
| 4 (35–40s) | "Comment 'DM' and I'll send you my three best CTA templates." | Avatar. Text: "Comment DM" |

## Using AI to draft scripts

Combine this lesson with your prompt skills:

> Write a 40-second vertical video script for [audience] about [topic]. Written for the ear: short sentences, contractions, no acronyms without phonetic spelling. Target 90–100 words. Format as a table with Scene, Voiceover, Visual. Split into 4 scenes. Hook in under 8 words. Do not add claims beyond these facts: [facts].

Then **read it aloud yourself** before generating. If you stumble, the AI voice will sound off too.

## Worked example: fixing a script line

**Before:** "Utilizing our cutting-edge, AI-driven, multi-channel CRM (customer relationship management) platform, SMEs can streamline their omnichannel engagement workflows."

**After:** "Here's the problem. Your customers message you on WhatsApp, Instagram and email. And you're losing track. Our CRM puts every message in one inbox."

Shorter, spoken, concrete, and much easier for an AI voice to deliver naturally.

## Hands-on: direction cues for expressive models

Expressive models such as ElevenLabs Eleven v3 accept inline audio tags in square brackets. Use them like a director's margin notes, sparingly:

```text
[warm] Most Reels get views. [short pause] Very few get DMs.
Here's why. [conversational] Your caption ends with "link in bio" - and that's a dead end.
[emphatic] Give people one word to comment instead.
```

Rules of thumb: one cue per sentence at most, test every cue with your chosen voice (a cue that suits one voice can sound odd on another), and keep a clean version without tags for models that read brackets literally. For stable long-form narration, rewriting the sentence is usually more reliable than a tag.

## A script prompt that produces producible scenes

```text
You are a scriptwriter for AI-voiced explainer videos.
Audience: {who}. Platform: {platform, aspect ratio}. Length: {seconds} s (~{words} words at 150 wpm).
Goal: {one action the viewer should take}.
Facts you may use (do not add others): {bullet facts}.
Write a table with columns: Scene | Seconds | Voiceover (spoken, short sentences) | On-screen text (max 6 words) | Visual (what the editor shows).
Rules: hook in under 8 words; one idea per scene; spell acronyms phonetically; numbers written as spoken;
no claims beyond the facts; end with one CTA. Then list 3 words the voice may mispronounce.
```

The last line is a small trick: asking the model to predict mispronunciations gives you a head start on the pronunciation list.

## Second worked example: a KSA real-estate developer

A Riyadh developer needs a 45-second Arabic and English teaser for a new community. The first AI draft is full of superlatives ("the most luxurious", "unmatched") that the legal team will not approve and that are hard to voice naturally. The rewrite uses concrete facts: distance to the metro, number of parks, handover quarter. Each scene has one fact and one visual. The Arabic version is written separately by a native writer from the same fact list rather than translated line by line, because sentence rhythm differs between languages. Both scripts are read aloud before any generation.

## Pitfalls

- Pasting blog copy straight into a voice tool.
- Scripts too long for the target length, which forces rushed speed settings.
- Forgetting to test pronunciation of local names such as place names, dishes and people.

## Video lecture: Scripting for AI voices and avatars

Lecture coming soon · 13 chapters · about 8 minutes. Read the full transcript below.

1. Scripting for AI voice and avatars
2. Why scripting matters most
3. Write for the ear
4. Practical rules
5. Sizing scripts
6. Short-form structure
7. The scene table
8. Example 1: fixing one line
9. Example 2: Riyadh developer
10. Watch me do it
11. Finishing touches
12. Common mistakes
13. Recap and try this now

## Lecture transcript

### Scripting for AI voice and avatars

Read this sentence out loud: utilizing our cutting-edge, AI-driven, multi-channel customer relationship management platform, small businesses can streamline their omnichannel engagement workflows. Did you run out of breath? So will an AI voice. It will read every word exactly as written, with none of the rescue a human narrator would give it. In this lecture you'll learn to write for the ear, size scripts to time, use a scene-based format that keeps voice and visuals aligned, and use AI to draft scripts you can actually produce.

### Why scripting matters most

Why is scripting the highest-leverage skill in AI media? Because everything downstream follows the script. The voice performs exactly what's on the page. The editor cuts to the scene breaks you wrote. The captions repeat your words. A weak script can't be rescued by a better voice or a fancier model. But a strong script makes even a plain library voice and simple visuals feel professional. If you have one hour to improve a video, spend it on the script.

### Write for the ear

Here's the core idea: write for the ear, not the eye. Think of the difference between a recipe card and a cooking show. A recipe card can use abbreviations, tables and long ingredient lists, because you can re-read it. A cooking show host says one thing at a time, repeats the key step, and tells you what to watch for. Your listener can't scroll back. So short sentences, one idea each. Conversational words: use, not utilize. Contractions. And front-load the point, because viewers decide in the first seconds.

### Practical rules

Now the practical rules. Write numbers the way you want them spoken. Two thousand five hundred rupees reads differently from R S two comma five hundred. Spell acronyms and brand names as they should sound. Use punctuation as direction: commas for breath, full stops for beats, a new paragraph for a longer pause. Avoid tongue-twisters and dense lists, because three items read aloud is plenty. And when you want emphasis, rephrase rather than rely on formatting. Here's the key part, is more reliable than bold text a voice can't see.

### Sizing scripts

Next, sizing. Comfortable narration runs at roughly one hundred thirty to one hundred sixty words per minute, depending on language, voice and style. So a fifteen-second video holds about thirty-five to forty words. Thirty seconds, about sixty-five to eighty. A minute, one hundred thirty to one hundred sixty. And leave room for pauses and visuals. The most common failure isn't bad writing. It's a ninety-second script crammed into a sixty-second slot, which forces a rushed speed setting and makes every voice sound anxious.

### Short-form structure

Then structure. For short-form video, use four beats. A hook in the first three seconds: a problem, a surprising statement, or a promise the video will keep. A setup until about ten seconds: why this matters to this viewer. The value, two or three concrete points, each with an example. And a call to action in the last five seconds, just one. If you're working with avatars, split scripts into scenes of ten to twenty seconds, and cut to b-roll, screen recordings or text between them. That keeps energy up and hides avatar limitations.

### The scene table

The format that holds it all together is the scene table. Columns for scene, seconds, voiceover, on-screen text and visual. Why a table? Because it forces you to decide what the viewer sees for every line the voice says. It keeps the voice producer, the editor and the reviewer looking at the same thing. And it maps directly onto per-scene audio files, which makes fixes cheap. If a line has no visual, that's a warning. If a visual has no line, it's probably b-roll, which is fine, as long as you planned it.

### Example 1: fixing one line

First example, a simple fix. The before line is the jargon sentence from the start. The after version is this. Here's the problem. Your customers message you on WhatsApp, Instagram and email. And you're losing track. Our system puts every message in one inbox. Four short sentences. Concrete channels the listener recognizes. A problem, then a promise. Same meaning, but now an AI voice can deliver it naturally, and a viewer can follow it without rewinding.

### Example 2: Riyadh developer

Second example, a business case. A Riyadh real-estate developer needs a forty-five-second Arabic and English teaser. The first AI draft is full of superlatives like most luxurious and unmatched. Legal won't approve them, and they're hard to voice without sounding like a parody. The rewrite uses concrete facts: minutes to the metro, number of parks, the handover quarter. One fact and one visual per scene. And the Arabic version is written by a native writer from the same fact list, not translated line by line, because sentence rhythm differs between languages.

### Watch me do it

Watch me do it. I paste the script prompt from the lesson into my assistant and fill the slots. Audience: small clinic owners. Platform: Instagram Reels, nine by sixteen. Length: forty seconds, about a hundred words. Goal: comment the word demo. Facts: three bullet points I've verified. I ask for a table with scene, seconds, voiceover, on-screen text and visual. The model returns five scenes. Then I read every voiceover line out loud with a timer running. Scene three runs long and I stumble on it, so I cut it into two shorter sentences.

### Finishing touches

One more trick from that prompt. The last line asks the model to list three words the voice might mispronounce. It flags the clinic software name, a Pakistani city, and an abbreviation. I add those to my pronunciation list before generating a single second of audio. And if I'm using an expressive model like Eleven v3, I might add one audio tag, like warm at the start, and test it. One cue per sentence at most, and I keep a clean version without tags for models that read brackets literally.

### Common mistakes

Common mistakes. Pasting blog copy straight into a voice tool. It was written for the eye. Scripts too long for the slot, which forces rushed speed. Forgetting to test pronunciation of local names like places, dishes and people. Letting the AI add claims beyond your facts. And skipping the read-aloud. Here's a simple test: if you stumble, the AI voice will sound off too. Your own mouth is the fastest quality check you have.

### Recap and try this now

Recap. Write for the ear: short sentences, conversational words, the point up front. Size scripts at about one hundred thirty to one hundred sixty words a minute and leave room to breathe. Use the scene table so every line has a visual. Draft with AI inside a prompt that restricts it to verified facts. And read everything out loud before you generate. Your voice is the cheapest test equipment you own. Try this now: take one paragraph from your website, rewrite it as a thirty-second scene table, read it aloud with a timer, and fix every line you stumble on.

## Key takeaways

- Write for the ear: short sentences, conversational words, a front-loaded hook.
- Size scripts at roughly 130–160 spoken words per minute, leaving room for pauses.
- Split avatar scripts into short scenes with b-roll and on-screen text between them.
- Use a two-column Scene / Voiceover / Visual format and read scripts aloud before generating.

## Try it

Rewrite a paragraph from your website or blog as a 30-second spoken script in the two-column format, read it aloud, and fix every line you stumble on.

- [Previous: The text-to-video landscape in 2026: Veo, Runway, Kling, Luma and life after Sora](https://optimizeall.com/learn/ai-video-and-voice-production/text-to-video-landscape-2026)
- [Next: Shot lists, storyboards and visual consistency](https://optimizeall.com/learn/ai-video-and-voice-production/shot-lists-and-storyboards)
- [All lessons of AI Video & Voice Production: ElevenLabs, Veo, Runway and More](https://optimizeall.com/learn/ai-video-and-voice-production)
