---
title: "Video understanding: timestamps, audio and frames"
description: "Two ways models \"watch\" video Video understanding means asking a model questions about a video: what happens, when, who says what, and what is shown on…"
url: https://optimizeall.com/learn/multimodal-and-reasoning-models/video-understanding
updated: 2026-10-05
---

Multimodal & Reasoning Models in Practice · Vision: documents, charts, screenshots and UI · lesson 4 of 17 · 12 min

# Video understanding: timestamps, audio and frames

## Two ways models "watch" video

Video understanding means asking a model questions about a video: what happens, when, who says what, and what is shown on screen. There are two technical approaches:

1. **Native video input.** Some models (notably Google's Gemini models) accept video files or supported video URLs directly. Internally the model samples frames at a set rate and, where available, processes the audio track too, so it can answer questions that combine speech and visuals, with timestamps.
2. **Frame sampling plus transcription.** For models that accept images but not video, you extract frames yourself (for example one every few seconds, or at scene changes), transcribe the audio separately, and send the frames with timestamps and the transcript.

The mental model: **a video is a sequence of sampled pictures plus an audio track.** Anything that happens between sampled frames may be missed. Fast actions, small on-screen text and brief moments need higher sampling rates or targeted clips.

## What video understanding is good for

- **Content operations:** chaptering long videos, generating titles, descriptions and highlight timestamps for YouTube or podcasts.
- **Compliance and brand checks:** does a creator's sponsored video include the required disclosure in the first seconds, and is the product shown accurately?
- **Training and support:** turning screen recordings into step-by-step guides.
- **Research:** summarising competitor ads or webinar recordings.
- **Quality assurance:** checking that a generated or edited video matches the storyboard.

## Hands-on: native video with the Gemini API

```python
import os, time
from google import genai

client = genai.Client()  # reads GEMINI_API_KEY
video = client.files.upload(file="ads/creator-spot-2026-09.mp4")
while video.state.name == "PROCESSING":        # large files are processed asynchronously
    time.sleep(5)
    video = client.files.get(name=video.name)

prompt = """You are reviewing a sponsored creator video for a skincare brand.
1. Give timestamps (mm:ss) for: first product appearance, first spoken
   mention of the brand, and any on-screen or spoken ad disclosure
   (e.g. "#ad", "paid partnership", "sponsored").
2. Is a clear disclosure present within the first 10 seconds? Yes/No, with evidence.
3. List any product claims made (quote them with timestamps).
If something is not visible or audible, say "not found"; do not guess."""

resp = client.models.generate_content(model=os.environ["GEMINI_MODEL"], contents=[video, prompt])
print(resp.text)
```

Google also offers the newer Interactions API as its primary interface for Gemini models and agents; the prompting principles are identical. Check the video-understanding documentation for supported formats, length limits, frame-rate settings and `media_resolution` options, and for how YouTube URLs can be passed.

## Hands-on: frame sampling for image-only models

```python
# pip install opencv-python
import cv2, base64

def sample_frames(path, every_seconds=3, max_frames=40):
    cap = cv2.VideoCapture(path)
    fps = cap.get(cv2.CAP_PROP_FPS) or 25
    step, frames, i = int(fps * every_seconds), [], 0
    while len(frames) < max_frames:
        cap.set(cv2.CAP_PROP_POS_FRAMES, i)
        ok, frame = cap.read()
        if not ok:
            break
        ok, jpg = cv2.imencode(".jpg", frame, [cv2.IMWRITE_JPEG_QUALITY, 80])
        frames.append((i / fps, base64.b64encode(jpg.tobytes()).decode()))
        i += step
    cap.release()
    return frames  # list of (seconds, base64 jpeg)
```

Send each frame as an image part preceded by a text label such as "Frame at 00:12", add the transcript with timestamps, and ask your question. Mind the cost: forty frames is forty images.

## Prompting tips for video

- **Ask for timestamps** in a fixed format so answers are checkable.
- **Scope the question:** "between 02:00 and 04:30" beats "the whole video".
- **Separate audio and visual claims:** "what is said" versus "what is shown".
- **Use permission to abstain:** "not found" rather than a guess.

## Worked example: a creator agency in London

An agency managing sponsored content for creators in the UK and the Gulf checks every draft video before it goes live. The model returns timestamps for disclosures and claims; a human opens each timestamp and confirms. Videos lacking a clear, early disclosure (as expected by UK advertising rules and by platforms' paid-partnership policies) go back to the creator. Review time falls, and the agency keeps a record of every check.

## Pitfalls

- Trusting a "no disclosure found" result without checking; low sampling rates can miss a brief caption.
- Sending hour-long videos when a short clip answers the question.
- Forgetting that faces and voices in video are personal data; follow consent and retention rules.

## How to measure success

Build a labelled set of 20 to 30 videos with known timestamps for key events. Measure timestamp accuracy (within a tolerance, such as two seconds), event recall and false "found" claims. Re-test when you change sampling rate, resolution or model.

## Video lecture: Video understanding: timestamps, audio and frames

Lecture coming soon · 11 chapters · about 7 minutes. Read the full transcript below.

1. Video understanding
2. The mental model
3. Use cases
4. Hands-on: native video (Gemini)
5. Hands-on: frame sampling
6. Prompting tips
7. Worked example: sponsored video checks
8. Example 1: a recorded team meeting
9. Example 2: branch training videos (illustrative)
10. Pitfalls and evaluation
11. Recap

## Lecture transcript

### Video understanding

A creator agency reviews dozens of sponsored videos every week, checking that each one discloses the partnership early and makes no unsupported product claims. Watching them all takes hours. A model can do the first pass in minutes, if you understand how models actually watch video. In this lecture you will learn the two ways models process video, what video understanding is good for, hands-on code for native video and for frame sampling, prompting tips, and how to evaluate the results.

### The mental model

The mental model: a video is a sequence of sampled pictures plus an audio track. Some models, notably Google's Gemini models, accept video directly. Internally they sample frames at a set rate and, where available, process the audio too, so they can answer questions combining speech and visuals, with timestamps. For models that accept images but not video, you do the sampling yourself: extract frames every few seconds or at scene changes, transcribe the audio separately, and send both. Either way, anything that happens between sampled frames may be missed. Fast actions, brief captions and small on-screen text need higher sampling or targeted clips.

### Use cases

Video understanding shines in content operations: chaptering long videos, and drafting titles, descriptions and highlight timestamps. In compliance and brand checks: does a sponsored video disclose the partnership early, and is the product shown accurately? In training and support: turning screen recordings into step-by-step guides. In research: summarising competitor ads or webinar recordings. And in quality assurance: checking that a generated or edited video matches its storyboard.

### Hands-on: native video (Gemini)

Here is the native approach with the Gemini API. You upload the video through the Files API, wait while it processes, then send the video and a prompt together. The prompt, for a sponsored skincare video, asks for mm colon ss timestamps for the first product appearance, the first spoken brand mention, and any on-screen or spoken disclosure, like hashtag ad or paid partnership. Then: is a clear disclosure present within the first ten seconds, yes or no, with evidence? Then list product claims with quotes and timestamps. And if something is not visible or audible, say not found, and do not guess. Google's newer Interactions API offers the same capabilities; the prompting principles are identical.

### Hands-on: frame sampling

For image-only models, the lesson includes a short frame sampler. It opens the video with OpenCV, reads the frame rate, grabs a frame every few seconds up to a maximum, encodes each as a compressed JPEG, and returns a list of timestamps and images. You then send each frame as an image part with a label like frame at zero zero twelve, add the timestamped transcript, and ask your question. Mind the cost: forty frames is forty images, so sample only as densely as the question requires.

### Prompting tips

Four prompting tips make video answers checkable. Ask for timestamps in a fixed format. Scope the question to a time range, like between two minutes and four thirty, rather than the whole video. Separate audio claims from visual claims, what is said versus what is shown. And always give permission to abstain with not found. Those four habits turn a vague summary into something a reviewer can verify in seconds by jumping to each timestamp.

### Worked example: sponsored video checks

Back to the London agency, which manages sponsored content for creators in the UK and the Gulf. Every draft video gets a model pass that returns timestamps for disclosures and claims. A human opens each timestamp and confirms. Videos without a clear, early disclosure, as expected under UK advertising rules and platforms' paid-partnership policies, go back to the creator. Review time falls, and the agency keeps a record of every check. But notice the human step. A no disclosure found result is a strong hint, not proof, because a brief caption can fall between sampled frames.

### Example 1: a recorded team meeting

A simple worked example. You recorded a forty-minute team meeting on video and want the moments when decisions were made. Instead of asking, summarise this video, ask: list every decision made, with the mm colon ss timestamp and who made it; separate what was said from what was shown on the shared screen; if no decision was made in a section, skip it. The model returns six decisions with timestamps. You jump to each one, confirm four, and correct one where the speaker was only proposing an idea. Your minutes take ten minutes instead of forty. The timestamps are what make it checkable.

### Example 2: branch training videos (illustrative)

Now a business scenario, with illustrative numbers. A fast-food chain in the UAE with about sixty branches receives short training videos made by each branch for a new food-safety procedure. Head office needs to confirm each video shows five required steps in order, including glove changes and temperature checks. Watching sixty videos takes a trainer two full days. They use native video understanding: for each video, list the time each of the five steps appears, whether the order is correct, and any step not found. The model processes all sixty in under an hour. It flags nine videos. The trainer reviews only those nine, plus five random passes as a spot-check, and confirms seven real issues and two missed moments the model could not see because the camera angle hid the thermometer. Total review time: about three hours instead of two days. Illustrative figures.

### Pitfalls and evaluation

Pitfalls to avoid: trusting a not-found result without checking; sending hour-long videos when a short clip answers the question; and forgetting that faces and voices in video are personal data, so follow consent and retention rules. To evaluate, build a labelled set of twenty to thirty videos with known timestamps for key events. Measure timestamp accuracy within a tolerance, like two seconds, event recall, and false found claims. Re-test whenever you change the sampling rate, resolution setting or model.

### Recap

To recap. Models watch video as sampled frames plus audio, so brief events can be missed. Use native video input where available, or sample frames and transcribe yourself. Ask for fixed-format timestamps, scope to time ranges, separate audio and visual, and allow not found. Verify negative findings, and evaluate on a labelled set. Try this now: label five real videos with the timestamps of three key events, run a video-understanding prompt, and count how many timestamps land within two seconds of the truth. Next module: audio and speech.

## Key takeaways

- Models watch video as sampled frames plus audio; events between samples can be missed.
- Use native video input where available, or sample frames and transcribe audio yourself.
- Ask for fixed-format timestamps, scope questions to time ranges and allow 'not found'.
- Evaluate timestamp accuracy and event recall on a labelled set before trusting results.

## Try it

Label five real videos with the timestamps of three key events, run a video-understanding prompt, and measure how many timestamps fall within two seconds of the truth.

- [Previous: Hands-on: sending images and PDFs through model APIs](https://optimizeall.com/learn/multimodal-and-reasoning-models/multimodal-api-inputs-images-and-pdfs)
- [Next: Speech-to-text: transcription that holds up](https://optimizeall.com/learn/multimodal-and-reasoning-models/speech-to-text)
- [All lessons of Multimodal & Reasoning Models in Practice](https://optimizeall.com/learn/multimodal-and-reasoning-models)
