Multimodal & Reasoning Models in PracticeImage and video generation concepts · Lesson 9 of 17
Video generation: concepts, limits and workflows
Video lecture
Video generation: concepts, limits and workflows
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Video generation
A consumer electronics start-up made a twenty-second teaser for a fraction of the cost of a shoot. The city at dawn and the abstract light trails were generated. The product itself only appeared in real footage. Nothing in it misrepresented what customers would get. In this lecture you will learn what today's video generators can do, how to direct a shot including its audio, a production workflow that actually works, how to call a video API, when generated video is the wrong choice, and the ethics and rules around synthetic video.
0:40 Why video generation matters
Why does this matter? Because video is the most engaging format on almost every platform, and the most expensive to produce. Generated video can fill the gaps: the establishing shot you cannot afford to film, the mood clip for a concept pitch, the b-roll that makes a talking-head video watchable. But it also carries the highest risk of misleading people. So the goal of this lecture is simple: use generation where it adds value, and real footage where honesty demands it.
1:15 What's possible now
Video generation models create short clips from text, images or other videos. Conceptually they extend image generation across time, producing frames consistent with each other and with plausible motion. The leading models now generate synchronised audio, including dialogue, sound effects and ambience. They accept reference images for characters and products, and support image-to-video and first-and-last-frame control, which makes storyboarding far more practical. Google's Veo models are available through the Gemini API and Google's creative tools, and Runway, Kling, Luma and others offer their own platforms and APIs.
1:53 Plan for change
The market also shows how quickly things change. OpenAI discontinued the Sora web and app experiences in April twenty twenty-six, and removed the Sora two video generation API on the twenty-fourth of September twenty twenty-six. If your workflow depends on a single video provider, keep your prompts, storyboards and assets portable, so you can switch without starting over. Treat any one tool as replaceable.
2:21 A shot as a mini script
Direct each shot like a mini script. Give it a number and duration, and the aspect ratio. Visual: close-up of hands pouring karak chai from a steel pot into a glass, steam rising, warm morning light in a small Dubai café. Camera: slow push-in, shallow depth of field. Audio: the gentle clink of glass and soft street ambience, no music. Dialogue: none. Style: realistic, warm colour grade, consistent with shots one and two. Specify audio explicitly. If you leave it out, some models add music or speech you did not want. And keep one clear action per shot.
3:04 Production workflow
Here is a workflow that works. Script and storyboard first, deciding shots and durations before generating anything. Generate key frames as images, get approval, then animate them with image-to-video for better control and consistency. Generate a few takes per shot and choose the best. Edit conventionally: cut, colour, sound, captions and brand elements in an editor. Review product accuracy, representation, cultural fit and legal claims. Then label and document according to platform rules.
3:36 Calling a video API
Video APIs are asynchronous, because generation is slow. With the Gemini API's Veo models, you call generate videos with a model name and prompt, which returns an operation. You poll that operation every few seconds until it is done, then download the generated video file and save it. Check the current model name and parameters in the docs. Budget for iteration, and record prompts and settings for every shot you use. A useful metric is cost per usable second: total generation spend and time, divided by the seconds that made the final cut.
4:16 Good fits and poor fits
Know the limits. Faces, clothing, products and text can change between frames or shots. Physics and causality can break, with objects passing through each other or hands deforming. Complex interactions between several people are hard. Precise control over brand assets, choreography or timing can take many attempts or conventional production. So use generated video for mood and b-roll, concept visualisation and pitches, stylised social content, and localisation with consent. Avoid it for accurate product demonstrations, testimonials, and news-like or documentary content where authenticity is expected.
4:53 Ethics and rules
Ethics and law matter more with video. Never generate realistic video of real people without consent, and never to deceive. Many jurisdictions are introducing rules on synthetic media, especially around elections, intimate imagery and fraud. Advertising regulators expect ads not to mislead, so generated demonstrations that exaggerate performance can breach the rules. Platforms increasingly require labels for realistic AI content. And provenance matters: content credentials and watermarks help audiences and platforms identify synthetic media, so preserve them where your tools support it.
5:29 Example 1: a five-second coffee clip
A simple worked example. You want a five-second clip of steam rising from a coffee cup for an Instagram story. First prompt: coffee, cosy, morning. The clip has music you did not want, the cup changes shape halfway, and the camera wanders. Second prompt, as a mini script: close-up of a white ceramic cup on a wooden table, steam rising, soft window light, static camera, audio: quiet café ambience, no music, realistic. One action, one camera instruction, explicit audio. The second clip is calm and usable. Specific direction beats mood words.
6:09 Example 2: a desert camp promo (illustrative)
Now a business scenario, with illustrative numbers. A tourism startup in Muscat needs a thirty-second promotional video for a desert camp. A traditional shoot would need crew travel and several days. They storyboard eight shots. Four are real footage filmed on a phone by staff: the actual tents, the dining area, the fire pit and guests with consent. Four are generated: dunes at dawn, a star-filled sky timelapse, camels on a ridge and abstract sand textures, all generic scenery that makes no claims about the camp. They generate three takes per generated shot, choose the best, and edit everything together with captions and music. Generation costs a small fraction of a shoot. Crucially, nothing generated shows the camp's facilities, so the ad cannot mislead. It is labelled per platform rules. Illustrative figures.
7:06 Recap
To recap. Video generators now produce short clips with audio, references and frame control, but products change, so keep assets portable. Direct each shot like a mini script with explicit audio and one clear action. Storyboard first, animate approved key frames, generate several takes, and edit conventionally. Use generated video for mood and concepts, show real products for real claims, and follow consent, disclosure and advertising rules. Try this now: storyboard a fifteen-second clip in four shots, and mark which you would generate, which you would film, and why. Next module: reasoning models.
What text-to-video systems do
Video generation models create short clips from text prompts, images or other videos. Conceptually, they extend image generation across time: generating frames that are consistent with each other and with plausible motion. Many use diffusion-based techniques over compressed video representations, and some can generate synchronised audio. Quality, clip length, resolution and control have advanced rapidly, and the landscape of tools changes frequently.
Separate from fully generative video, there are avatar and presenter tools (a digital presenter speaking your script) and AI editing tools (auto-captions, cutting, reframing, dubbing). Our AI video and voice course covers specific tools; here we focus on concepts that apply across them.
Directing a generated shot
Video prompts combine image-prompt elements with motion and camera language:
- Subject and action: what happens, in one clear beat.
- Setting: place, time of day, atmosphere.
- Camera: shot size (wide, medium, close-up), movement (static, slow push-in, pan, tracking), angle.
- Style: realistic, animated, cinematic, documentary.
- Duration and pacing within the tool's limits.
A slow push-in on a ceramic coffee cup on a café counter in the early
morning; steam rising; soft window light; a barista's hands place a
pastry beside it. Realistic, shallow depth of field, calm pace, 16:9.One clear action per shot works better than a complex sequence. Build longer stories by generating multiple shots and editing them together.
Common limitations
- Temporal consistency: faces, clothing, products and text can change between frames or shots.
- Physics and causality: objects may pass through each other, liquids behave oddly, hands deform.
- Complex actions and interactions between several people are hard.
- Precise control: matching exact brand assets, choreography or timing often requires many attempts, reference images, or conventional production.
- Cost and time: generation can be expensive and slow relative to images; iteration budgets matter.
A practical production workflow
- Script and storyboard first. Decide the story, shots and durations before generating anything.
- Generate key frames as images, get approval, then animate them (image-to-video) for better control and consistency.
- Generate several takes per shot, choose the best.
- Edit conventionally: cut, colour, sound, captions, brand elements in an editor.
- Review: accuracy of product depiction, representation, cultural fit, legal claims.
- Label and document in line with platform rules.
When to use generated video
Good fits:
- Mood and b-roll shots (atmosphere, abstract backgrounds, transitions).
- Concept visualisation and pitching before a real shoot.
- Quick social content where stylised visuals are acceptable.
- Localisation: dubbing and translated versions (with consent from speakers).
Poor fits:
- Accurate product demonstrations (show the real product).
- Testimonials or anything implying real people's experiences.
- News-like or documentary content where authenticity is expected.
Ethics, law and trust
- Deepfakes: never generate realistic video of real people without consent, and never to deceive. Many jurisdictions have introduced or are introducing rules on synthetic media, particularly around elections, intimate imagery and fraud.
- Advertising standards: advertising regulators in many countries expect ads not to mislead; generated demonstrations that exaggerate performance can breach these rules.
- Disclosure: platforms increasingly require labels for realistic AI-generated content. Follow each platform's current policy.
- Provenance: content credentials and watermarking help audiences and platforms identify synthetic media; preserve them where your tools support it.
Worked example: a product launch teaser
A consumer electronics start-up makes a 20-second teaser. Generated shots provide atmosphere: a city at dawn, abstract light trails. The product itself appears only in real footage shot on a phone gimbal. Captions and logo are added in the editor. The final teaser is labelled according to the platforms' AI content rules. Total cost is a fraction of a full shoot, and nothing in it misrepresents the product.
What changed recently
Leading video models now generate clips with synchronised audio (dialogue, sound effects, ambience), accept reference images for characters and products, and support image-to-video and first-and-last-frame control, which makes storyboarding far more practical. Google's Veo models are available through the Gemini API and Google's creative tools; Runway, Kling, Luma and others offer their own platforms and APIs.
The landscape also shows how quickly products can change. OpenAI discontinued the Sora web and app experiences in April 2026 and removed the Sora 2 video generation API on 24 September 2026. If a workflow depends on one video provider, keep prompts, storyboards and assets portable so you can switch.
Prompting video with audio
Treat each shot as a mini script:
SHOT 3 (6 seconds, 16:9)
Visual: Close-up of hands pouring karak chai from a steel pot into a glass,
steam rising, warm morning light in a small Dubai café.
Camera: slow push-in, shallow depth of field.
Audio: gentle clink of glass, soft street ambience, no music.
Dialogue: none.
Style: realistic, warm colour grade, consistent with shots 1-2.Specify audio explicitly. If you leave it out, some models add music or speech you did not want.
Hands-on: generating a clip via an API (asynchronous pattern)
Video generation is slow, so APIs return a job you poll. With the Gemini API's Veo models the pattern looks like this (check the current model name and parameters in the docs):
import os, time
from google import genai
client = genai.Client() # reads GEMINI_API_KEY
operation = client.models.generate_videos(
model=os.environ["VEO_MODEL"],
prompt=open("shots/shot3.txt", encoding="utf-8").read(),
)
while not operation.done:
time.sleep(10)
operation = client.operations.get(operation)
video = operation.response.generated_videos[0]
client.files.download(file=video.video)
video.video.save("shot3.mp4")Budget for iteration: generate a few takes per shot, keep the best, and record prompts and settings for every shot you use.
Going further
Track your own "cost per usable second": total generation spend and time divided by seconds that made the final cut. It is a more honest metric than the price per generation, and it helps you decide which shots to generate versus film.
Key takeaways
- Video generators extend image generation across time; avatar and editing tools are related but distinct.
- Direct shots with subject and one action, setting, camera, style and pacing; build stories from multiple shots.
- Limitations include temporal consistency, physics, complex interactions, precise control and cost.
- Use generated video for mood, concepts and b-roll; show real products; never deceive, and follow disclosure rules.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Storyboard a 15-second social clip in four shots. Mark which shots you would generate, which you would film, and why.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.