Multimodal & Reasoning Models in PracticeImage and video generation concepts · Lesson 9 of 17

Video generation: concepts, limits and workflows

Article · 11 min · 8 min lecture

Video lecture

Video generation: concepts, limits and workflows

12 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 12

Video generation

  • What generators do now
  • Directing shots with audio
  • A production workflow
  • When not to use it

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What text-to-video systems do

Video generation models create short clips from text prompts, images or other videos. Conceptually, they extend image generation across time: generating frames that are consistent with each other and with plausible motion. Many use diffusion-based techniques over compressed video representations, and some can generate synchronised audio. Quality, clip length, resolution and control have advanced rapidly, and the landscape of tools changes frequently.

Separate from fully generative video, there are avatar and presenter tools (a digital presenter speaking your script) and AI editing tools (auto-captions, cutting, reframing, dubbing). Our AI video and voice course covers specific tools; here we focus on concepts that apply across them.

Directing a generated shot

Video prompts combine image-prompt elements with motion and camera language:

  1. Subject and action: what happens, in one clear beat.
  2. Setting: place, time of day, atmosphere.
  3. Camera: shot size (wide, medium, close-up), movement (static, slow push-in, pan, tracking), angle.
  4. Style: realistic, animated, cinematic, documentary.
  5. Duration and pacing within the tool's limits.
A slow push-in on a ceramic coffee cup on a café counter in the early
morning; steam rising; soft window light; a barista's hands place a
pastry beside it. Realistic, shallow depth of field, calm pace, 16:9.

One clear action per shot works better than a complex sequence. Build longer stories by generating multiple shots and editing them together.

Common limitations

  • Temporal consistency: faces, clothing, products and text can change between frames or shots.
  • Physics and causality: objects may pass through each other, liquids behave oddly, hands deform.
  • Complex actions and interactions between several people are hard.
  • Precise control: matching exact brand assets, choreography or timing often requires many attempts, reference images, or conventional production.
  • Cost and time: generation can be expensive and slow relative to images; iteration budgets matter.

A practical production workflow

  1. Script and storyboard first. Decide the story, shots and durations before generating anything.
  2. Generate key frames as images, get approval, then animate them (image-to-video) for better control and consistency.
  3. Generate several takes per shot, choose the best.
  4. Edit conventionally: cut, colour, sound, captions, brand elements in an editor.
  5. Review: accuracy of product depiction, representation, cultural fit, legal claims.
  6. Label and document in line with platform rules.

When to use generated video

Good fits:

  • Mood and b-roll shots (atmosphere, abstract backgrounds, transitions).
  • Concept visualisation and pitching before a real shoot.
  • Quick social content where stylised visuals are acceptable.
  • Localisation: dubbing and translated versions (with consent from speakers).

Poor fits:

  • Accurate product demonstrations (show the real product).
  • Testimonials or anything implying real people's experiences.
  • News-like or documentary content where authenticity is expected.

Ethics, law and trust

  • Deepfakes: never generate realistic video of real people without consent, and never to deceive. Many jurisdictions have introduced or are introducing rules on synthetic media, particularly around elections, intimate imagery and fraud.
  • Advertising standards: advertising regulators in many countries expect ads not to mislead; generated demonstrations that exaggerate performance can breach these rules.
  • Disclosure: platforms increasingly require labels for realistic AI-generated content. Follow each platform's current policy.
  • Provenance: content credentials and watermarking help audiences and platforms identify synthetic media; preserve them where your tools support it.

Worked example: a product launch teaser

A consumer electronics start-up makes a 20-second teaser. Generated shots provide atmosphere: a city at dawn, abstract light trails. The product itself appears only in real footage shot on a phone gimbal. Captions and logo are added in the editor. The final teaser is labelled according to the platforms' AI content rules. Total cost is a fraction of a full shoot, and nothing in it misrepresents the product.

What changed recently

Leading video models now generate clips with synchronised audio (dialogue, sound effects, ambience), accept reference images for characters and products, and support image-to-video and first-and-last-frame control, which makes storyboarding far more practical. Google's Veo models are available through the Gemini API and Google's creative tools; Runway, Kling, Luma and others offer their own platforms and APIs.

The landscape also shows how quickly products can change. OpenAI discontinued the Sora web and app experiences in April 2026 and removed the Sora 2 video generation API on 24 September 2026. If a workflow depends on one video provider, keep prompts, storyboards and assets portable so you can switch.

Prompting video with audio

Treat each shot as a mini script:

SHOT 3 (6 seconds, 16:9)
Visual: Close-up of hands pouring karak chai from a steel pot into a glass,
steam rising, warm morning light in a small Dubai café.
Camera: slow push-in, shallow depth of field.
Audio: gentle clink of glass, soft street ambience, no music.
Dialogue: none.
Style: realistic, warm colour grade, consistent with shots 1-2.

Specify audio explicitly. If you leave it out, some models add music or speech you did not want.

Hands-on: generating a clip via an API (asynchronous pattern)

Video generation is slow, so APIs return a job you poll. With the Gemini API's Veo models the pattern looks like this (check the current model name and parameters in the docs):

import os, time
from google import genai

client = genai.Client()  # reads GEMINI_API_KEY
operation = client.models.generate_videos(
    model=os.environ["VEO_MODEL"],
    prompt=open("shots/shot3.txt", encoding="utf-8").read(),
)
while not operation.done:
    time.sleep(10)
    operation = client.operations.get(operation)

video = operation.response.generated_videos[0]
client.files.download(file=video.video)
video.video.save("shot3.mp4")

Budget for iteration: generate a few takes per shot, keep the best, and record prompts and settings for every shot you use.

Going further

Track your own "cost per usable second": total generation spend and time divided by seconds that made the final cut. It is a more honest metric than the price per generation, and it helps you decide which shots to generate versus film.

Key takeaways

  • Video generators extend image generation across time; avatar and editing tools are related but distinct.
  • Direct shots with subject and one action, setting, camera, style and pacing; build stories from multiple shots.
  • Limitations include temporal consistency, physics, complex interactions, precise control and cost.
  • Use generated video for mood, concepts and b-roll; show real products; never deceive, and follow disclosure rules.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which prompt structure works best for a generated video shot?
  2. Which use is the poorest fit for generated video?
  3. Why generate key frames as images before animating them?

Put it into practice

Storyboard a 15-second social clip in four shots. Mark which shots you would generate, which you would film, and why.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.