---
title: "The text-to-video landscape in 2026: Veo, Runway, Kling…"
description: "The landscape moves monthly: learn to evaluate, not to memorize Text-to-video is the fastest-moving part of this course. In the twelve months to…"
url: https://optimizeall.com/learn/ai-video-and-voice-production/text-to-video-landscape-2026
updated: 2026-10-05
---

AI Video & Voice Production: ElevenLabs, Veo, Runway and More · Generated video: avatars, b-roll and text-to-video models · lesson 7 of 16 · 9 min

# The text-to-video landscape in 2026: Veo, Runway, Kling, Luma and life after Sora

## The landscape moves monthly: learn to evaluate, not to memorize

Text-to-video is the fastest-moving part of this course. In the twelve months to September 2026, new model generations shipped from several vendors, native audio became common, and one of the best-known products disappeared. The durable skill is not knowing today's leaderboard. It is running a fair, quick evaluation for **your** use case and choosing a tool you can legally and reliably use.

## Where things stand (September 2026)

Treat the table as a snapshot; verify availability and terms on each vendor's official site before a project.

| Model family / product | Access routes | Notable strengths (vendor-described) | Watch-outs |
|---|---|---|---|
| **Google Veo 3.1** (Veo 3.1, Veo 3.1 Fast, Veo 3.1 Lite) | Gemini app and Flow for creators; Gemini API (Google AI Studio); Vertex AI, now part of Google Cloud's Gemini Enterprise Agent Platform | Native audio (dialogue, effects), image-to-video, reference images for consistency, first/last-frame control, video extension; higher-resolution output options | Outputs carry Google's invisible **SynthID** watermark; preview models and quotas change; person-generation rules vary by region |
| **Runway Gen-4.5** | Runway web app; Runway API | Prompt fidelity, realistic motion and physics; strong editing tools around generation | Credits-based pricing; check commercial terms per plan |
| **Kling 3.0** (Kuaishou) | Kling web app; API and third-party platforms | Multi-shot sequences, native audio in several languages, high-resolution output | Data-location and terms review for enterprise use |
| **Luma Ray3 family** (Dream Machine) | Luma web app; API | Keyframe control, video modification of existing footage, HDR options | Features differ by plan and model version |
| **OpenAI Sora** | **Discontinued**: the Sora app and web experience closed on April 26, 2026, and the Sora 2 models and Videos API were removed from the OpenAI API on September 24, 2026 | n/a | Do not build new workflows on it; migrate existing ones |

Two lessons from the Sora shutdown are worth writing on a sticky note. First, **any generator can disappear or change terms**, so keep your prompts, reference images and edit projects portable. Second, **export and archive** everything you need: assets inside a closed product can be deleted when it shuts down.

## How to choose: the five-test evaluation

Run the same five prompts on every candidate, with the same aspect ratio and duration, and score blind (hide which tool made which clip).

```text
T1 Concept b-roll      "Five glowing folders merge into one above a desk at dusk, slow push-in"
T2 Product-in-context  Image-to-video from YOUR product photo: "slow orbit, soft studio light"
T3 Character consistency  Same generated presenter-free character across two shots (reference image)
T4 Motion and physics  "Chai poured into a glass, steam rising, close-up"
T5 Your market         A culturally specific scene you actually need (e.g., Karachi street at dusk,
                       Riyadh office, London bus stop) - check for stereotypes and errors
Score 1-5: prompt adherence | artifacts | motion realism | consistency | audio (if used)
Also record: cost per usable second, time to first result, license terms, C2PA/SynthID behavior
```

**Cost per usable second** is the metric that surprises people. If a cheaper tool needs six generations for one usable clip and a pricier one needs two, the pricier one may be cheaper in practice.

## Hands-on: generate a clip with Veo through the Gemini API

This uses Google's official `google-genai` Python SDK. Video generation is a long-running operation: you start it, poll it, then download the result.

```python
# pip install google-genai
import os, time
from google import genai
from google.genai import types

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

operation = client.models.generate_videos(
    model="veo-3.1-generate-preview",          # check the current model id in the docs
    source=types.GenerateVideosSource(
        prompt=("Close-up of hands pouring karak chai into a small glass on a wooden table, "
                "steam rising, soft morning light, slow push-in, shallow depth of field. "
                "Quiet room tone, gentle pouring sound. No text, no faces."),
    ),
    config=types.GenerateVideosConfig(aspect_ratio="9:16", number_of_videos=1),
)

while not operation.done:            # generation takes a while
    time.sleep(15)
    operation = client.operations.get(operation)

if operation.error:
    raise RuntimeError(operation.error)

clip = operation.response.generated_videos[0]
client.files.download(file=clip.video)       # fetches the bytes into clip.video
clip.video.save("chai_opener.mp4")
print("saved chai_opener.mp4")
```

Notes: model ids with `preview` can change or be retired, quotas and pricing are per second of video, and your organization's cloud terms apply when you use the enterprise route. Keep the prompt, model id and date in your production log.

## Consistency techniques that work across tools

- **Reference images** (image-to-video, "ingredients"/reference features): feed the same character, product or style frame into every shot.
- **First and last frame control**: define where a shot starts and ends so cuts line up.
- **Shot extension**: extend a good clip instead of regenerating a new one.
- **Fixed seeds** where offered, so small prompt edits change less.
- **Edit, don't regenerate**: color-grade and crop in the editor to unify shots from different tools.

## Worked example: an agency's quarterly tool review

A Dubai content agency produces b-roll for twelve clients. Every quarter it runs the five-test evaluation on three candidate tools, scores blind with two editors, and computes cost per usable second. In the last review (illustrative), one tool won concept b-roll, another won image-to-video product shots, and none handled culturally specific Gulf office scenes without errors, so those stay filmed. The agency keeps two subscriptions rather than one, and writes the results into a one-page "tool policy" that account managers show clients.

## Outlook (reasoned, not predicted)

Expect three trends to continue: longer and multi-shot generation, native audio as standard, and more provenance marking because regulation (such as the EU AI Act's transparency rules applying from August 2026) and platform labeling push vendors that way. Plan workflows that can swap models without rewriting your whole process.

## Pitfalls

- Choosing from a vendor's showcase reel instead of your own five tests.
- Ignoring cost per usable second.
- Building a workflow on a single product with no export plan.
- Stripping watermarks or content credentials to avoid labels.

## Video lecture: The text-to-video landscape in 2026: Veo, Runway, Kling, Luma and life after Sora

Lecture coming soon · 13 chapters · about 8 minutes. Read the full transcript below.

1. The text-to-video landscape, 2026
2. Why evaluate, not memorize
3. Snapshot: September 2026
4. Lessons from Sora
5. The five tests
6. Scoring and practical facts
7. Example 1: a cooking creator
8. Example 2: a Dubai agency
9. Watch me do it, part 1
10. Watch me do it, part 2
11. Common mistakes
12. Recap
13. Try this now

## Lecture transcript

### The text-to-video landscape, 2026

In March twenty twenty-six, OpenAI announced it was shutting down Sora, one of the most famous video generators in the world. The app closed in April, and the models were removed from the API in September. If your workflow depended on it, you had a problem. In this lecture you'll get a clear snapshot of the text-to-video landscape as of September twenty twenty-six. More importantly, you'll learn a five-test evaluation you can rerun every quarter, and you'll generate a clip through an API yourself.

### Why evaluate, not memorize

Why focus on evaluation rather than a leaderboard? Because this is the fastest-moving part of AI media. New model generations ship every few months, native audio went from rare to common, and a product can disappear. If you memorize today's rankings, you'll be wrong by the next quarter. If you know how to test tools fairly against your own needs, you'll make good decisions no matter what launches next. That's a skill clients and employers will pay for.

### Snapshot: September 2026

Here's the snapshot. Google's Veo three point one family, including Fast and Lite variants, is available in the Gemini app and Flow for creators, and in the Gemini API and Vertex AI, which Google now calls the Gemini Enterprise Agent Platform. It generates native audio and supports reference images and first and last frame control. Runway Gen four point five is available in Runway's app and API. Kling three point zero from Kuaishou emphasizes multi-shot sequences. And Luma's Ray three family offers keyframe control and video modification. Verify each on the vendor's site before a project.

### Lessons from Sora

And the Sora story teaches two lessons. First, any generator can disappear or change its terms, so keep your prompts, reference images and edit projects portable. Think of your tools like rented equipment in a film studio. You can rent the best camera in town, but your script, storyboard and footage should always live in your own cupboard. Second, export and archive everything you need. When a product shuts down, assets left inside it can be deleted. Portable inputs and archived outputs protect you.

### The five tests

Now the five-test evaluation. Run the same five prompts on every candidate, same aspect ratio, same duration. Test one, concept b-roll. Test two, image-to-video from your actual product photo. Test three, consistency of a character or object across two shots. Test four, motion and physics, like pouring liquid. And test five, a scene from your market, like a Karachi street at dusk or a Riyadh office, checking for errors and stereotypes. Then score blind, meaning the reviewers don't know which tool made which clip.

### Scoring and practical facts

Score each clip from one to five on prompt adherence, artifacts, motion realism, consistency, and audio if you use it. Then record four practical facts: time to first result, license terms, how the tool marks its output, and cost per usable second. That last one surprises people. If a cheaper tool needs six generations to get one usable clip, and a pricier one needs two, the pricier one may be cheaper in practice. Track the usable seconds, not the sticker price.

### Example 1: a cooking creator

First example, a simple one. A creator wants vertical b-roll for cooking Shorts. She runs tests one and four on two tools she already has access to, with the same chai-pouring prompt. Tool A looks gorgeous but the liquid pours upward in two out of three takes. Tool B is slightly softer but pours correctly every time. She picks tool B for motion shots and keeps tool A for static mood shots. Two tests, twenty minutes, and a decision grounded in her own content.

### Example 2: a Dubai agency

Second example, a business case with illustrative results. A Dubai agency produces b-roll for twelve clients. Every quarter, two editors run the full five tests on three candidate tools and score blind. In the latest round, one tool won concept b-roll, another won product image-to-video, and none handled culturally specific Gulf office scenes without errors, so those are still filmed. The agency keeps two subscriptions and writes a one-page tool policy that account managers share with clients. Clients like knowing the choice was tested, not guessed.

### Watch me do it, part 1

Watch me do it through an API. I use Google's official Python SDK. I create a client with my key from an environment variable. I call generate videos with the Veo model id from the docs, and a director-style prompt that also describes the sound: quiet room tone, a gentle pour. I set the aspect ratio to nine by sixteen. Video generation is a long-running operation, so the call returns immediately with an operation, and I poll it every fifteen seconds until it's done.

### Watch me do it, part 2

When it's done, I check for an error first, because a safety filter or quota limit can stop a job. If it's clean, I download the generated file and save it as chai opener dot mp4. Then I do the part people skip. I open my production log and write the model id, the full prompt, the date, and a note that Google marks Veo output with an invisible SynthID watermark. Preview model ids change, so that log is how I'll reproduce or explain this clip in six months.

### Common mistakes

Common mistakes. Choosing a tool from its showcase reel instead of your own tests. Showcases are cherry-picked by definition. Ignoring cost per usable second. Building your whole workflow on one product with no export plan, which is exactly what hurt Sora users. And stripping watermarks or content credentials to avoid platform labels. That undermines trust, can break platform rules, and runs against the direction of regulation like the EU AI Act's transparency rules, which apply from August twenty twenty-six.

### Recap

Recap. The landscape as of September twenty twenty-six includes Veo three point one, Runway Gen four point five, Kling three, and the Luma Ray three family, and Sora has been discontinued. Don't memorize rankings. Run the five tests, score blind, and compare cost per usable second. Keep your inputs portable and archive your outputs. Log model ids, prompts and dates. Looking ahead, and this is outlook rather than prediction, expect longer multi-shot clips, native audio as standard, and more provenance marking.

### Try this now

Try this now. Pick two video tools you can access today. Run tests one, two and five from the lesson with identical settings. Ask a colleague to score the clips blind, and calculate cost per usable second. Write the result as a three-line tool policy: which tool for which job, and what you'll still film for real. Put a date on it and a reminder to rerun the tests next quarter. Then, if you're comfortable with a little code, run the Veo script from the lesson with your own prompt, and log the model id, the prompt and the date. You'll have both a decision and a working API pipeline by the end of the day.

## Key takeaways

- Evaluate tools with your own five tests and blind scoring rather than vendor showcase reels.
- As of September 2026, major options include Veo 3.1, Runway Gen-4.5, Kling 3.0 and the Luma Ray3 family; Sora was discontinued.
- Compare cost per usable second, not sticker price.
- Keep prompts and references portable, archive outputs, and log model ids, prompts and dates.

## Try it

Run tests T1, T2 and T5 on two tools you can access, have a colleague score blind, compute cost per usable second and write a three-line tool policy.

- [Previous: AI b-roll, text-to-video and generated visuals](https://optimizeall.com/learn/ai-video-and-voice-production/ai-broll-and-text-to-video)
- [Next: Scripting for AI voices and avatars](https://optimizeall.com/learn/ai-video-and-voice-production/scripting-for-ai-voice-and-avatars)
- [All lessons of AI Video & Voice Production: ElevenLabs, Veo, Runway and More](https://optimizeall.com/learn/ai-video-and-voice-production)
