---
title: "How vision-language models see | Optimize All Academy"
description: "From pixels to tokens Many current AI models are multimodal : they accept images (and sometimes audio or video) alongside text. In a typical…"
url: https://optimizeall.com/learn/multimodal-and-reasoning-models/how-vision-models-see
updated: 2026-10-05
---

Multimodal & Reasoning Models in Practice · Vision: documents, charts, screenshots and UI · lesson 1 of 17 · 11 min

# How vision-language models see

## From pixels to tokens

Many current AI models are **multimodal**: they accept images (and sometimes audio or video) alongside text. In a typical vision-language model, an image is split into patches, each converted into a numerical representation by a vision encoder, and these become "visual tokens" that the language model processes together with your text. The model does not see an image the way you do; it sees a compressed representation whose detail depends on image resolution, how the provider resizes images, and how many visual tokens are used.

This mental model explains most practical behaviour:

- **Resolution matters.** Tiny text in a large screenshot may be lost when the image is downscaled. Crop to the region that matters, or send a higher-resolution image if your provider supports it.
- **Images cost tokens.** Larger images generally consume more input tokens, affecting cost and latency. Check your provider's documentation for how images are sized and billed.
- **Layout is understood, imperfectly.** Models can follow tables, forms and diagrams, but complex layouts (merged cells, multi-column text, rotated labels) are where errors cluster.

## What vision models are good at

- Reading printed text in reasonably clear images and documents.
- Describing scenes, products and UI screens.
- Extracting structured data from forms, receipts and invoices.
- Interpreting the gist of charts and diagrams.
- Comparing two images (before and after, design versions).
- Answering questions that combine an image with domain knowledge ("Which step of this setup guide has the user reached?").

## Where they struggle

- **Precise measurement and counting:** exact pixel positions, counting many small objects, reading precise values off a chart axis.
- **Dense or degraded input:** handwriting, low-light photos, faint stamps, heavily compressed scans.
- **Spatial reasoning:** left/right relationships in complex scenes, small differences between near-identical images.
- **Hallucinated detail:** confidently "reading" text that is too blurry to read, or inferring content that is not present.

Capabilities are improving quickly, and they vary between models. Treat these as areas to test rather than fixed limits.

## Prompting with images

The fundamentals from text prompting still apply, plus some image-specific habits:

```text
You are checking a supplier invoice image for our accounts team.

1. Extract: supplier name, invoice number, invoice date, currency,
   subtotal, tax, total.
2. If any field is unreadable or absent, write "UNREADABLE" or "ABSENT";
   do not guess.
3. Report any visual issues (cropping, blur, handwriting) in "notes".

Return JSON matching this schema: {...}
```

- **Say what the image is** and why you are sending it.
- **Ask for "unreadable" rather than guesses.** This is the vision equivalent of permission to say "I don't know".
- **Place images before the question** when the provider allows ordering, similar to long documents in text prompts; check provider guidance.
- **Label multiple images** ("Image 1: current homepage; Image 2: proposed redesign").
- **Crop and split.** For a dense page, send regions separately or at higher resolution.

## Worked example: expense receipts

A finance team processes receipts from staff in the UK, UAE and Pakistan. First attempt: send photos, ask for totals. Problems: faded thermal receipts misread, currency symbols confused, totals occasionally invented for unreadable receipts. Improvements:

1. Require ISO currency codes and a `readability` field.
2. Ask for line items and total separately, then check in code that items sum to the total within tolerance.
3. Route unreadable or inconsistent receipts to a human.
4. Ask staff to photograph receipts flat, in good light (a process fix, often the biggest win).

## Evaluating vision tasks

Build a test set of real images, including the hard ones: crumpled receipts, screenshots on dark mode, charts with dual axes. For each, record the correct extraction. Measure field-level accuracy, rate of correctly flagged unreadable fields, and rate of confident errors (wrong values not flagged). Confident errors matter most, because they slip through.

## Resolution controls you can actually use

Providers now expose resolution and detail controls, because image tokens drive cost and latency:

- **Gemini** offers a `media_resolution` setting (low, medium, high and, on some models, ultra-high) that sets how many tokens an image, PDF page or video frame receives; on Gemini 3 models it can be set per media item.
- **OpenAI** image inputs accept a detail level (for example low or high) that trades fidelity for tokens.
- **Claude** resizes very large images automatically; sending images near the recommended size avoids wasted upload time without changing quality, and cropping remains the most effective way to give small details more pixels.

Exact token counts and limits change by model, so check the vision documentation for your model. The durable rule: **spend resolution where the detail is.** A receipt total needs pixels; a scene description does not.

## Hands-on: an image extraction you can trust

```python
import base64, json, os, pathlib
import anthropic

client = anthropic.Anthropic()
MODEL = os.environ.get("CLAUDE_MODEL", "claude-opus-5")

SCHEMA = {
    "type": "object",
    "properties": {
        "supplier": {"type": "string"},
        "invoice_date": {"type": "string", "description": "YYYY-MM-DD, or UNREADABLE / ABSENT"},
        "currency": {"type": "string", "description": "ISO code such as GBP, AED, PKR, SAR, or UNREADABLE"},
        "total": {"type": ["number", "null"], "description": "null if unreadable or absent"},
        "readability": {"type": "string", "enum": ["clear", "partly_readable", "unreadable"]},
        "notes": {"type": "string", "description": "Cropping, blur, handwriting, glare"},
    },
    "required": ["supplier", "invoice_date", "currency", "total", "readability", "notes"],
    "additionalProperties": False,
}

def extract(path: str) -> dict:
    data = base64.standard_b64encode(pathlib.Path(path).read_bytes()).decode()
    media_type = "image/png" if path.lower().endswith(".png") else "image/jpeg"
    resp = client.messages.create(
        model=MODEL, max_tokens=1024,
        output_config={"format": {"type": "json_schema", "schema": SCHEMA}},
        messages=[{"role": "user", "content": [
            {"type": "image", "source": {"type": "base64", "media_type": media_type, "data": data}},
            {"type": "text", "text": "This is a supplier invoice photo for our accounts team. "
             "Extract the fields. Never guess: use UNREADABLE or ABSENT, or null for the total."},
        ]}],
    )
    return json.loads(next(b.text for b in resp.content if b.type == "text"))

print(extract("receipts/2026-09-12-dubai.jpg"))
```

The image comes **before** the question, the schema forces an explicit readability judgement, and the total can be null, so the model is never pushed to invent a value.

## Measuring confident errors

For every test image, compare each field with the truth and classify: correct, correctly flagged (UNREADABLE when it truly was), confident error (a wrong value not flagged), or over-flagged (UNREADABLE when it was readable). Confident errors are the metric to drive towards zero; over-flagging costs human time but not trust.

## Going further

Vision can be combined with traditional tools. Dedicated OCR engines can extract text with character positions; the language model can then interpret it. For high-volume document processing, compare a pure vision-model approach against OCR plus a text model on your documents: accuracy, cost and latency may favour either, depending on your data.

## Video lecture: How vision-language models see

Lecture coming soon · 11 chapters · about 8 minutes. Read the full transcript below.

1. How vision-language models see
2. Pixels → patches → visual tokens
3. Strengths vs struggles
4. Spend resolution where detail is
5. Prompting with images
6. Hands-on: a trustworthy extraction
7. Worked example: expense receipts
8. Evaluating vision
9. Example 1: splitting a restaurant bill
10. Example 2: proof-of-delivery photos (illustrative)
11. Recap

## Lecture transcript

### How vision-language models see

A finance team sends a model a photo of a faded thermal receipt and asks for the total. It answers instantly and confidently: forty-two pounds fifty. The real total was twenty-four fifty. The model did not lie on purpose. It could not see well enough, and nobody gave it permission to say so. In this lecture you will learn how vision-language models turn pixels into something they can reason about, what they are good and bad at, how to spend resolution where it matters, and how to write image prompts that produce trustworthy extractions.

### Pixels → patches → visual tokens

Here is the mental model. In a typical vision-language model, an image is split into patches. A vision encoder converts each patch into numbers, and these become visual tokens that the language model processes alongside your text. The model does not see the image the way you do. It sees a compressed representation, and how much detail survives depends on the image's resolution, how the provider resizes it, and how many visual tokens it gets. Three consequences follow. Resolution matters, so tiny text in a big screenshot may be lost. Images cost tokens, so bigger images cost more and take longer. And layout is understood, but imperfectly.

### Strengths vs struggles

What are vision models good at? Reading printed text in reasonably clear images, describing scenes, products and app screens, extracting structured data from forms, receipts and invoices, getting the gist of charts, comparing two images, and combining an image with domain knowledge, like which step of this setup guide has the user reached. Where do they struggle? Precise measurement and counting many small objects. Dense or degraded input, like handwriting, low light and faint stamps. Fine spatial reasoning. And, most dangerously, hallucinated detail: confidently reading text that is too blurry to read.

### Spend resolution where detail is

Providers now give you resolution controls, because image tokens drive cost and latency. Gemini has a media resolution setting, from low up to high and, on some models, ultra-high, that sets the tokens per image, PDF page or video frame, and on Gemini three models you can set it per item. OpenAI image inputs take a detail level. Claude resizes very large images automatically. Exact numbers change by model, so check the docs. The durable rule is: spend resolution where the detail is. A receipt total needs pixels. A scene description does not. And cropping is still the most effective way to give small details more pixels.

### Prompting with images

Now prompting. Everything from text prompting still applies, plus some image habits. Say what the image is and why you are sending it: a supplier invoice for our accounts team. Put the image before the question. Label multiple images: image one is the current homepage, image two the proposed redesign. Crop and split dense pages. And, most importantly, give the model permission to say it cannot read something. Ask for UNREADABLE or ABSENT instead of guesses. That is the vision equivalent of permission to say I don't know.

### Hands-on: a trustworthy extraction

The lesson's hands-on code puts this together. It reads a receipt photo, encodes it, and sends it to the Claude API with a JSON schema. The schema has fields for supplier, invoice date, currency as an ISO code, the total, a readability judgement of clear, partly readable or unreadable, and notes about blur or cropping. The total is allowed to be null. So the model is never forced to invent a value just to fill a field. The same pattern works with any provider that supports structured outputs.

### Worked example: expense receipts

Here is a worked example. A finance team processes receipts from staff in the UK, the UAE and Pakistan. First attempt: send photos, ask for totals. Problems: faded thermal receipts misread, currency symbols confused, and totals invented for unreadable receipts. Improvements: require ISO currency codes and a readability field; ask for line items and the total separately, then check in code that items sum to the total; route unreadable or inconsistent receipts to a person. And the biggest win was a process fix: ask staff to photograph receipts flat, in good light.

### Evaluating vision

How do you evaluate a vision task? Build a test set of real images, including the hard ones: crumpled receipts, dark-mode screenshots, charts with dual axes. Record the correct values. Then classify every field: correct, correctly flagged as unreadable, confident error, or over-flagged. Confident errors, wrong values that were not flagged, are the metric to drive toward zero, because they slip straight through. Over-flagging costs human time, but not trust. And compare approaches: for high-volume documents, a dedicated OCR engine plus a text model can beat a pure vision model on cost or accuracy, depending on your data.

### Example 1: splitting a restaurant bill

Let us work through a simple example. You photograph a restaurant bill to split it with friends, and ask a model to list each item and price. The photo is taken at an angle, and the bottom line is in shadow. Version one of your prompt: list the items and total. The model lists everything, including a total it partly guessed. Version two: crop the photo to just the bill, take it flat, and ask, list each item and price; if any price is unclear, write unclear; then give the total only if every line is readable. Now the model marks one price as unclear and withholds the total. You check that one line yourself. Better input, plus permission to be uncertain, turned a guess into a trustworthy answer.

### Example 2: proof-of-delivery photos (illustrative)

Now a business scenario, with illustrative numbers. A courier company in Karachi processes about twelve thousand proof-of-delivery photos a month: parcels on doorsteps with handwritten signatures and address labels. They want to auto-check that the label's tracking number matches the job. A first test on three hundred photos finds about ninety-one percent matched, but eleven confident errors, where the model read a wrong digit without flagging it. The team changes three things. They ask drivers to photograph the label close up, as well as the doorstep. They send the close-up at higher resolution and the doorstep shot at low. And they require a readability field plus an exact-match check in code. On the next three hundred photos, confident errors drop to one, and unreadable labels go to a small review queue. Illustrative figures, but the combination of better capture, targeted resolution and code checks is the pattern.

### Recap

To recap. Vision models turn images into visual tokens, so resolution, resizing and token budgets shape what they can see. They are strong at clear text, description and extraction, weaker at counting, degraded input and fine detail. Spend resolution where the detail is, crop, label images, put them before the question, and allow UNREADABLE. Measure confident errors most closely. Try this now: collect ten real images you would process, write a prompt with an UNREADABLE option, and score field-level accuracy and confident errors. Next: documents, charts, screenshots and user interfaces.

## Video transcript

In this course we're going beyond text, into models that can see, hear, speak, generate images and video, reason at length, and even operate a computer. Let's begin with vision.

Here's the mental model. When you send an image to a vision-language model, it doesn't see it the way you do. The image is resized, cut into patches, and turned into what you can think of as visual tokens that the language model reads alongside your text. That explains a lot. If the text in your screenshot is tiny, it may get lost when the image is scaled down. So crop to what matters. And bigger images usually cost more tokens, so there's a trade-off between detail and cost.

Vision models are genuinely good at reading clear documents, extracting fields from forms and receipts, describing screens and products, and getting the gist of a chart. Where they struggle is precision: counting lots of small objects, reading exact values off a chart axis, handwriting, blurry photos, and fine spatial details. And like text models, they can hallucinate, confidently reading text that isn't really legible.

So give them an escape hatch. Ask them to write "unreadable" instead of guessing. Label your images, explain why you're sending them, and check totals and key values in code.

Finally, evaluate on your real images, especially the ugly ones: faded receipts, dark-mode screenshots, charts with two axes. And track the most dangerous kind of error: the confident one that nobody flagged.

## Key takeaways

- Vision models turn images into visual tokens; resolution, resizing and token budget shape what they can see.
- They are strong at reading clear text, describing, extracting and summarising charts; weaker at precise measurement, counting, degraded input and fine spatial detail.
- Prompt with context, labelled images, cropping and an explicit 'unreadable' option instead of guesses.
- Evaluate on real, hard images and track confident errors most closely.

## Try it

Collect 10 real images you would process (receipts, screenshots, charts). Write a prompt with an UNREADABLE option and score field-level accuracy and confident errors.

- [Next: Working with documents, charts, screenshots and UI](https://optimizeall.com/learn/multimodal-and-reasoning-models/documents-charts-screenshots-ui)
- [All lessons of Multimodal & Reasoning Models in Practice](https://optimizeall.com/learn/multimodal-and-reasoning-models)
