---
title: "AI b-roll, text-to-video and generated visuals"
description: "What text-to-video can do in 2026 Text-to-video and image-to-video models generate short clips from a prompt or a reference image. Quality has improved…"
url: https://optimizeall.com/learn/ai-video-and-voice-production/ai-broll-and-text-to-video
updated: 2026-10-05
---

AI Video & Voice Production: ElevenLabs, Veo, Runway and More · Generated video: avatars, b-roll and text-to-video models · lesson 6 of 16 · 9 min

# AI b-roll, text-to-video and generated visuals

## What text-to-video can do in 2026

Text-to-video and image-to-video models generate short clips from a prompt or a reference image. Quality has improved rapidly: many tools can produce cinematic shots, product-style visuals and stylized scenes, typically in short clips of a few seconds to tens of seconds that you then edit together. Many avatar and editing platforms now integrate these generators directly.

For marketers, the sweet spot is usually **b-roll**: supporting footage that illustrates what a narrator or avatar is saying, such as abstract concepts, establishing shots, product-in-context mood shots and transitions.

## Good uses

- **Illustrating concepts:** "a time-lapse of a city skyline shifting from night to day" for a productivity script.
- **Mood and atmosphere:** steaming chai on a rainy window, desert dunes at golden hour.
- **Stylized or animated explainers** where realism does not matter.
- **Background loops and transitions.**
- **Storyboarding and pre-visualization** before a real shoot.

## Risky or inappropriate uses

- **Showing your actual product performing** in a way it does not. Generated footage of a product that looks better, bigger or more effective than reality is misleading advertising.
- **Fake "real" events:** generated news-like footage, crowds at an event that did not happen, or before-and-after results.
- **Real people and places** that could be misrepresented, including recognizable individuals, landmarks in sensitive contexts or religious sites.
- **Imitating a specific artist's or brand's distinctive style**, which raises copyright and trademark concerns and can damage reputation.

A simple rule: **generated visuals may illustrate ideas; they must not stand in as evidence.**

## Prompting for video

Video prompts work best when they describe a shot the way a director would:

- **Subject:** what is in the frame.
- **Action:** what moves and how.
- **Setting:** location, time of day, weather.
- **Camera:** shot type (close-up, wide), movement (slow push-in, pan, handheld), lens feel.
- **Style and lighting:** photorealistic, 3D animation, film grain, soft daylight, neon.
- **Duration and aspect ratio:** vertical 9:16 for Reels, Shorts and TikTok; 16:9 for YouTube.

**Example prompt:**
"Close-up of hands pouring karak chai into a small glass on a wooden table, steam rising, soft morning light from a window, slow push-in, shallow depth of field, photorealistic, 9:16."

Generate several variations, pick the best and trim. Expect artifacts such as odd hands, flickering text and morphing objects, and check every frame used.

## Combining sources

A practical explainer might combine:

- Your real product footage (shot on a phone) for anything factual.
- An avatar or your own face for the presenter.
- AI b-roll for concepts and atmosphere.
- Licensed stock footage where you need real-world locations.
- Motion graphics for data and key points.

## Rights and licenses

- Check the tool's terms for **commercial use** of outputs on your plan.
- Copyright protection for purely AI-generated material is limited or uncertain in several jurisdictions, so you may not be able to stop others from copying it. Human creative contribution such as editing and arrangement can strengthen your position.
- Avoid prompts naming living artists, film studios or brands to copy their style.
- Keep a record of prompts and tools used for important commercial assets.

## Worked example

A Jeddah travel agency makes a 30-second Reel promoting a family Umrah package. They use real footage of partner hotels (licensed), generated b-roll only for abstract transitions such as a stylized map animation and a sunrise sky, a human presenter from the team, and on-screen prices checked against the package terms. They avoid generating footage of the holy sites themselves, both for accuracy and out of respect for religious sensitivities.

## Hands-on: a b-roll prompt card

Keep prompts structured so results are repeatable and reviewable:

```text
B-ROLL PROMPT CARD  (scene 04, 9:16, 6 s)
Purpose:     illustrate "your data lives in five places"
Subject:     five glowing folders floating above a wooden desk
Action:      folders drift slowly toward a laptop and merge into one
Setting:     home office at dusk, city lights through the window
Camera:      slow push-in, 35mm feel, shallow depth of field
Style:       photorealistic, soft warm key light, subtle film grain
Audio:       none (voice-over added in edit)  | or: "quiet room tone"
Negative:    no text, no logos, no people's faces, no hands
Reference:   brand color frame (image-to-video) if the tool supports it
Review:      artifacts? text morphing? anything that looks like product proof?
```

The **Negative** line matters: many models struggle with on-screen text, hands and faces, so ask for none and add text in the editor instead.

## Second worked example: a UK B2B explainer

A London cybersecurity start-up needs a 90-second explainer. Real screen recordings of the product show every factual claim. AI b-roll covers only abstract ideas: a padlock forming from data points, a stylized map with pulses of traffic. The team generates four variations per shot, keeps the best, and logs each prompt, model and date in the production record. One generated shot shows a logo-like shape on a server rack, so they cut it: if a viewer could read it as a real brand, it does not belong in the video.

## Pitfalls

- Letting generated b-roll dominate so the video feels hollow.
- Missing small artifacts that viewers screenshot and mock.
- Using AI visuals as product proof.

## Video lecture: AI b-roll, text-to-video and generated visuals

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

1. AI b-roll and text-to-video
2. Why it matters
3. The rule
4. Good uses versus risky uses
5. Prompt like a director
6. A practical source mix
7. Example 1: a tea brand opener
8. Example 2: London cybersecurity explainer
9. Watch me do it
10. Rights and records
11. Review routine
12. Common mistakes
13. Recap
14. Try this now

## Lecture transcript

### AI b-roll and text-to-video

Picture this. Your narrator says, your data lives in five different places. On screen, five glowing folders drift across a desk at dusk and merge into one. Nobody filmed that. A text-to-video model generated it from a paragraph. In this lecture you'll learn where generated b-roll genuinely helps, where it crosses a line, how to write prompts like a director, and how to combine generated shots with real footage so your video stays honest.

### Why it matters

Why does this matter? Because b-roll is where small teams used to lose the most time and money. Stock libraries rarely have the exact shot you need, and custom shoots are expensive. Generated b-roll closes that gap for concepts, moods and transitions. But the same tool can create footage that looks like evidence. A product that looks bigger than it is. A crowd at an event that never happened. That's where marketing turns into misleading advertising, and where platforms and regulators get involved.

### The rule

Here's the rule for the whole lesson. Generated visuals may illustrate ideas. They must never stand in as evidence. Think of generated b-roll like the illustrations in a textbook. A drawing of the water cycle helps you understand, but nobody thinks it's a photograph of an actual cloud. The moment a generated clip could be mistaken for proof, of how your product performs, of a real event, or of a real person doing something, it has left the illustration zone.

### Good uses versus risky uses

Where does generated b-roll work well? Illustrating abstract concepts, like a city skyline shifting from night to day for a productivity script. Mood and atmosphere, like steaming chai against a rainy window, or desert dunes at golden hour. Stylized or animated explainers where realism doesn't matter. Background loops and transitions. And storyboarding before a real shoot. Where is it risky? Showing your actual product performing, fake before and after results, news-like footage, recognizable people or sensitive places, and imitating a specific artist's or brand's style.

### Prompt like a director

Now, prompting. Write prompts the way a director briefs a camera crew. Subject: what's in the frame. Action: what moves, and how. Setting: location, time of day, weather. Camera: shot type, movement, lens feel. Style and lighting: photorealistic, three D animation, soft daylight, neon. Then duration and aspect ratio: vertical nine by sixteen for Reels, Shorts and TikTok, sixteen by nine for YouTube. And add a negative line. Many models still struggle with on-screen text, hands and faces, so ask for none of them and add text later in your editor.

### A practical source mix

Most good videos combine sources. Here's a practical mix for an explainer. Your real product footage, even shot on a phone, for anything factual. A presenter, whether that's you, a voice over slides, or an avatar. Generated b-roll for concepts and atmosphere. Licensed stock footage where you need real-world locations. And motion graphics for data and key points. Think of it like a meal with several dishes. Generated b-roll is the garnish and the sauce. It makes the plate look great, but it isn't the protein. The facts come from real footage.

### Example 1: a tea brand opener

First example, a simple one. A tea brand wants a six-second vertical opener. The prompt: close-up of hands pouring karak chai into a small glass on a wooden table, steam rising, soft morning light from a window, slow push-in, shallow depth of field, photorealistic, nine by sixteen. They generate four variations. One has a warped glass. One has six fingers. They keep the cleanest take, trim it to five seconds, and add the brand name as text in the editor, not in the prompt. Four minutes of work, zero reshoots.

### Example 2: London cybersecurity explainer

Second example, a business case. A London cybersecurity start-up needs a ninety-second explainer. Every factual claim is shown with real screen recordings of the product. Generated b-roll covers only abstract ideas: a padlock forming from data points, a stylized map with pulses of traffic. They generate four variations per shot and log each prompt, model and date in a production record. One generated shot shows a logo-like shape on a server rack. They cut it. If a viewer could read it as a real brand, it doesn't belong in the video.

### Watch me do it

Watch me do it. I open the prompt card template from the lesson. Purpose: illustrate your data lives in five places. Subject: five glowing folders above a wooden desk. Action: folders drift toward a laptop and merge. Setting: home office at dusk. Camera: slow push-in, shallow depth of field. Style: photorealistic, warm key light. Audio: none, because the voice-over comes later. Negative: no text, no logos, no faces, no hands. I paste it into the video tool, generate four takes, and review them frame by frame before I choose.

### Rights and records

Now the rights side, because it's easy to forget. Check the tool's terms for commercial use on your plan. In several countries, copyright protection for purely AI-generated material is limited or uncertain, so others might be able to reuse it. Your human contribution, like editing, arrangement and combining with real footage, can strengthen your position. Avoid prompts that name living artists, studios or brands. And keep a record of prompts, tools and dates for important commercial assets. If a client or platform asks how something was made, you'll have the answer.

### Review routine

And here's my review routine before anything generated goes into an edit. I scrub every clip at quarter speed, looking at five hot spots: hands, faces, text, edges of objects, and anything that moves fast. I check that nothing looks like a real brand, a recognizable person, or a sensitive place. I ask one question: could a reasonable viewer mistake this for proof of something? If the answer is yes, it either gets cut or gets a clear label. It takes a few minutes per video, and it's the cheapest insurance you'll ever buy.

### Common mistakes

Common mistakes. Letting generated b-roll dominate so the video feels hollow. Viewers sense when nothing on screen is real. Missing small artifacts that people screenshot and mock, like melting text or extra fingers. Using AI visuals as product proof. And forgetting disclosure when realistic generated scenes could be mistaken for real events. We'll cover platform labels properly in module five, but the habit starts here.

### Recap

Let's recap. Generated b-roll is strongest for concepts, mood, transitions and stylized scenes. It must never stand in as evidence. Prompt like a director with a negative line, generate several takes, and inspect every frame. Mix it with real footage for anything factual, log your prompts, and check commercial-use terms. Review every clip at quarter speed for hands, faces, text and brand-like shapes. And keep asking the one question that keeps you honest: could a reasonable viewer mistake this for proof? If so, cut it or label it.

### Try this now

Try this now. Pick your next video and find three moments where the narration describes an idea rather than a fact. Write a prompt card for each, generate them in any tool you have access to, and note every artifact you'd need to cut. Then add each prompt to a simple production log with the tool, the model and the date. Finally, place the best clip into your edit next to real footage and watch the sequence on a phone. If the generated shot feels like decoration that supports the narration, you've got the balance right. If it feels like the star of the show, trim it back.

## Key takeaways

- Text-to-video is strongest for b-roll: concepts, mood, transitions and stylized scenes.
- Generated visuals may illustrate ideas but must never serve as product proof or fake events.
- Prompt like a director: subject, action, setting, camera, style, duration and aspect ratio.
- Check commercial-use terms, avoid copying artists' styles and inspect every frame for artifacts.

## Try it

Write three director-style b-roll prompts for your next video, generate them in any tool you have access to, and note any artifacts you would need to cut.

- [Previous: AI avatars and talking-head videos](https://optimizeall.com/learn/ai-video-and-voice-production/avatars-and-talking-heads)
- [Next: The text-to-video landscape in 2026: Veo, Runway, Kling, Luma and life after Sora](https://optimizeall.com/learn/ai-video-and-voice-production/text-to-video-landscape-2026)
- [All lessons of AI Video & Voice Production: ElevenLabs, Veo, Runway and More](https://optimizeall.com/learn/ai-video-and-voice-production)
