AI Video & Voice Production: ElevenLabs, Veo, Runway and MoreGenerated video: avatars, b-roll and text-to-video models · Lesson 7 of 16

The text-to-video landscape in 2026: Veo, Runway, Kling, Luma and life after Sora

Article · 9 min · 8 min lecture

Video lecture

The text-to-video landscape in 2026: Veo, Runway, Kling, Luma and life after Sora

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

The text-to-video landscape, 2026

  • Snapshot of the main models
  • Lessons from the Sora shutdown
  • A five-test evaluation
  • Generating via API

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The landscape moves monthly: learn to evaluate, not to memorize

Text-to-video is the fastest-moving part of this course. In the twelve months to September 2026, new model generations shipped from several vendors, native audio became common, and one of the best-known products disappeared. The durable skill is not knowing today's leaderboard. It is running a fair, quick evaluation for your use case and choosing a tool you can legally and reliably use.

Where things stand (September 2026)

Treat the table as a snapshot; verify availability and terms on each vendor's official site before a project.

Model family / productAccess routesNotable strengths (vendor-described)Watch-outs
Google Veo 3.1 (Veo 3.1, Veo 3.1 Fast, Veo 3.1 Lite)Gemini app and Flow for creators; Gemini API (Google AI Studio); Vertex AI, now part of Google Cloud's Gemini Enterprise Agent PlatformNative audio (dialogue, effects), image-to-video, reference images for consistency, first/last-frame control, video extension; higher-resolution output optionsOutputs carry Google's invisible SynthID watermark; preview models and quotas change; person-generation rules vary by region
Runway Gen-4.5Runway web app; Runway APIPrompt fidelity, realistic motion and physics; strong editing tools around generationCredits-based pricing; check commercial terms per plan
Kling 3.0 (Kuaishou)Kling web app; API and third-party platformsMulti-shot sequences, native audio in several languages, high-resolution outputData-location and terms review for enterprise use
Luma Ray3 family (Dream Machine)Luma web app; APIKeyframe control, video modification of existing footage, HDR optionsFeatures differ by plan and model version
OpenAI SoraDiscontinued: the Sora app and web experience closed on April 26, 2026, and the Sora 2 models and Videos API were removed from the OpenAI API on September 24, 2026n/aDo not build new workflows on it; migrate existing ones

Two lessons from the Sora shutdown are worth writing on a sticky note. First, any generator can disappear or change terms, so keep your prompts, reference images and edit projects portable. Second, export and archive everything you need: assets inside a closed product can be deleted when it shuts down.

How to choose: the five-test evaluation

Run the same five prompts on every candidate, with the same aspect ratio and duration, and score blind (hide which tool made which clip).

T1 Concept b-roll      "Five glowing folders merge into one above a desk at dusk, slow push-in"
T2 Product-in-context  Image-to-video from YOUR product photo: "slow orbit, soft studio light"
T3 Character consistency  Same generated presenter-free character across two shots (reference image)
T4 Motion and physics  "Chai poured into a glass, steam rising, close-up"
T5 Your market         A culturally specific scene you actually need (e.g., Karachi street at dusk,
                       Riyadh office, London bus stop) - check for stereotypes and errors
Score 1-5: prompt adherence | artifacts | motion realism | consistency | audio (if used)
Also record: cost per usable second, time to first result, license terms, C2PA/SynthID behavior

Cost per usable second is the metric that surprises people. If a cheaper tool needs six generations for one usable clip and a pricier one needs two, the pricier one may be cheaper in practice.

Hands-on: generate a clip with Veo through the Gemini API

This uses Google's official google-genai Python SDK. Video generation is a long-running operation: you start it, poll it, then download the result.

# pip install google-genai
import os, time
from google import genai
from google.genai import types

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

operation = client.models.generate_videos(
    model="veo-3.1-generate-preview",          # check the current model id in the docs
    source=types.GenerateVideosSource(
        prompt=("Close-up of hands pouring karak chai into a small glass on a wooden table, "
                "steam rising, soft morning light, slow push-in, shallow depth of field. "
                "Quiet room tone, gentle pouring sound. No text, no faces."),
    ),
    config=types.GenerateVideosConfig(aspect_ratio="9:16", number_of_videos=1),
)

while not operation.done:            # generation takes a while
    time.sleep(15)
    operation = client.operations.get(operation)

if operation.error:
    raise RuntimeError(operation.error)

clip = operation.response.generated_videos[0]
client.files.download(file=clip.video)       # fetches the bytes into clip.video
clip.video.save("chai_opener.mp4")
print("saved chai_opener.mp4")

Notes: model ids with preview can change or be retired, quotas and pricing are per second of video, and your organization's cloud terms apply when you use the enterprise route. Keep the prompt, model id and date in your production log.

Consistency techniques that work across tools

  • Reference images (image-to-video, "ingredients"/reference features): feed the same character, product or style frame into every shot.
  • First and last frame control: define where a shot starts and ends so cuts line up.
  • Shot extension: extend a good clip instead of regenerating a new one.
  • Fixed seeds where offered, so small prompt edits change less.
  • Edit, don't regenerate: color-grade and crop in the editor to unify shots from different tools.

Worked example: an agency's quarterly tool review

A Dubai content agency produces b-roll for twelve clients. Every quarter it runs the five-test evaluation on three candidate tools, scores blind with two editors, and computes cost per usable second. In the last review (illustrative), one tool won concept b-roll, another won image-to-video product shots, and none handled culturally specific Gulf office scenes without errors, so those stay filmed. The agency keeps two subscriptions rather than one, and writes the results into a one-page "tool policy" that account managers show clients.

Outlook (reasoned, not predicted)

Expect three trends to continue: longer and multi-shot generation, native audio as standard, and more provenance marking because regulation (such as the EU AI Act's transparency rules applying from August 2026) and platform labeling push vendors that way. Plan workflows that can swap models without rewriting your whole process.

Pitfalls

  • Choosing from a vendor's showcase reel instead of your own five tests.
  • Ignoring cost per usable second.
  • Building a workflow on a single product with no export plan.
  • Stripping watermarks or content credentials to avoid labels.

Key takeaways

  • Evaluate tools with your own five tests and blind scoring rather than vendor showcase reels.
  • As of September 2026, major options include Veo 3.1, Runway Gen-4.5, Kling 3.0 and the Luma Ray3 family; Sora was discontinued.
  • Compare cost per usable second, not sticker price.
  • Keep prompts and references portable, archive outputs, and log model ids, prompts and dates.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Your team built its b-roll workflow entirely inside the Sora app. What is the most important lesson from its 2026 shutdown?
  2. Tool A costs less per generation but needs about six attempts per usable clip; Tool B costs more but needs about two. What should you compare?
  3. Why does the five-test evaluation include a scene from your own market?

Put it into practice

Run tests T1, T2 and T5 on two tools you can access, have a colleague score blind, compute cost per usable second and write a three-line tool policy.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.