Skip to content

Multimodal & Reasoning Models in Practice · Vision: documents, charts, screenshots and UI · lesson 3 of 17 · 13 min

Hands-on: sending images and PDFs through model APIs

One request, many kinds of content

Every major model API now accepts a list of content parts in a single message: text, images, PDFs and, for some models, audio and video. The mental model is simple: you are assembling a briefing pack. Each part is labelled, placed in a sensible order, and followed by a clear question.

The shapes differ by provider, and knowing them saves hours:

| Provider | Image part | Document part | Reuse large files | |---|---|---|---| | Claude (Messages API) | {"type": "image", "source": ...} (base64, URL or file ID) | {"type": "document", "source": ...} for PDFs and text, with optional citations | Files API | | OpenAI (Responses API) | {"type": "input_image", ...} | {"type": "input_file", ...} | Files API | | Gemini | Inline data part or uploaded file | PDF as a file part | Files API; media_resolution controls tokens |

Check each provider's current documentation for size limits, supported formats and page limits, which differ by model.

Ordering and labelling multiple inputs

  • Media before the question. Put images and documents first, then your instructions and question.
  • Label every item. "Image 1: current checkout page. Image 2: competitor checkout page." With documents, give a title.
  • Say what to do with each. "Compare Image 1 against the design spec in the PDF; ignore Image 2 except for layout ideas."
  • Keep IDs through outputs. Ask the model to refer to "Image 1" or "page 7" in its answer so results are traceable.

Hands-on: a mixed-media request in Claude and OpenAI

Claude: compare two screenshots against a PDF brand guide.

import base64, os, pathlib, anthropic
client = anthropic.Anthropic()

def b64(path):
    return base64.standard_b64encode(pathlib.Path(path).read_bytes()).decode()

content = [
    {"type": "text", "text": "Image 1: our current landing page (mobile)."},
    {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": b64("shots/current.png")}},
    {"type": "text", "text": "Image 2: proposed redesign (mobile)."},
    {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": b64("shots/proposed.png")}},
    {"type": "document", "title": "Brand guide 2026",
     "source": {"type": "base64", "media_type": "application/pdf", "data": b64("brand/guide-2026.pdf")}},
    {"type": "text", "text": "Check Image 2 against the brand guide: colours, type scale, logo "
     "clear space, button style. List each issue with the guide page it relates to, and say "
     "whether Image 1 had the same issue. Mark anything you cannot judge from a static image "
     "as 'needs testing'."},
]
try:
    resp = client.messages.create(model=os.environ.get("CLAUDE_MODEL", "claude-opus-5"),
                                  max_tokens=4096, messages=[{"role": "user", "content": content}])
    print("".join(b.text for b in resp.content if b.type == "text"))
except anthropic.BadRequestError as err:      # e.g. file too large or unsupported format
    print("Request rejected:", err.message)

OpenAI (Responses API): the same idea, with different part names.

import base64, os
from openai import OpenAI
client = OpenAI()

img = base64.b64encode(open("shots/proposed.png", "rb").read()).decode()
uploaded = client.files.create(file=open("brand/guide-2026.pdf", "rb"), purpose="user_data")

resp = client.responses.create(
    model=os.environ["OPENAI_MODEL"],
    input=[{"role": "user", "content": [
        {"type": "input_text", "text": "Image: proposed redesign (mobile). File: brand guide."},
        {"type": "input_image", "image_url": f"data:image/png;base64,{img}"},
        {"type": "input_file", "file_id": uploaded.id},
        {"type": "input_text", "text": "List brand-guide violations in the image, citing guide pages."},
    ]}],
)
print(resp.output_text)

Reusing files and controlling cost

Uploading the same 80-page PDF on every request wastes bandwidth and tokens. Use the providers' Files APIs to upload once and reference by ID, and combine that with prompt caching for repeated questions over the same document. For images, crop and downscale to what the task needs; on Gemini, choose media_resolution per item so a detailed invoice gets more tokens than a decorative photo.

Worked example: a property agency in Dubai

A property agency receives tenancy applications as a mix of phone photos (passports are excluded by policy), salary letters as PDFs and screenshots of bank transfers. Their pipeline sends each application as one labelled request: "Document 1: salary letter; Image 1: transfer screenshot", asks for a structured summary with a source label for every field, and routes any field marked UNREADABLE or inconsistent to an agent. Personal data is minimised before sending, retention is configured to policy, and the provider's data terms were reviewed before launch, in line with the UAE's data protection law.

Pitfalls

  • Mixing up which image is which because items were not labelled.
  • Sending full-resolution photos when a cropped region would do.
  • Forgetting request size limits and getting rejected requests in production; handle the error and split inputs.
  • Sending personal or sensitive documents without checking data terms and retention settings.

How to measure success

Track field-level accuracy against a labelled set, the rate of correctly flagged unreadable fields, tokens per request and request failures due to size limits. A good multimodal pipeline is accurate, traceable to source items, and predictable in cost.

Video lecture: Hands-on: sending images and PDFs through model APIs

Lecture coming soon · 11 chapters · about 7 minutes. Read the full transcript below.

  1. Sending images and PDFs through APIs
  2. Content parts by provider
  3. Four habits
  4. Hands-on: Claude mixed request
  5. Hands-on: OpenAI Responses
  6. Reuse and cost control
  7. Worked example: tenancy applications
  8. Example 1: flyer vs brand guide
  9. Example 2: motor insurance claims (illustrative)
  10. Pitfalls and metrics
  11. Recap

Lecture transcript

Sending images and PDFs through APIs

You want a model to compare two versions of your landing page against an eighty-page brand guide. That is three very different inputs in one question. In this lecture you will learn how model APIs accept mixed content, the part types for Claude, OpenAI and Gemini, how to order and label inputs so answers stay traceable, how to reuse large files and control cost, and how to handle the errors and privacy questions that come with real documents.

Content parts by provider

Every major API now accepts a list of content parts in one message: text, images, PDFs, and for some models audio and video. Think of it as assembling a briefing pack. Each item is labelled, placed in a sensible order, and followed by a clear question. The shapes differ by provider. Claude uses image blocks and document blocks, with sources that can be base64, a URL, or a file ID. OpenAI's Responses API uses input text, input image and input file parts. Gemini uses inline data or uploaded file parts, with media resolution controlling tokens. Check each provider's docs for size, format and page limits.

Four habits

Four habits keep mixed inputs reliable. Media before the question: images and documents first, then instructions. Label every item: image one is our current landing page on mobile; image two is the proposed redesign; and give documents a title. Say what to do with each: check image two against the brand guide, and only compare with image one for context. And ask for answers that refer to those labels or to page numbers, so every finding is traceable to a source.

Hands-on: Claude mixed request

Here is the Claude version from the lesson. It builds a content list: a text label, then the first screenshot as a base64 image block, another label, the second screenshot, then the brand guide as a document block with a title, and finally the instruction. Check image two against the brand guide for colours, type scale, logo clear space and button style; list each issue with the guide page; say whether image one had the same issue; and mark anything you cannot judge from a static image as needs testing. The call is wrapped in error handling, because oversized files or unsupported formats return a bad request error you should catch.

Hands-on: OpenAI Responses

The OpenAI version follows the same idea with different names. It encodes the screenshot as a data URL for an input image part, uploads the PDF once through the Files API with the user data purpose, and references it by file ID in an input file part. The instruction then asks for brand-guide violations with guide pages. Different syntax, same briefing-pack principle.

Reuse and cost control

Uploading the same eighty-page PDF with every question wastes bandwidth and tokens. Use the providers' Files APIs to upload once and reference by ID, and combine that with prompt caching for repeated questions over the same document. For images, crop and downscale to what the task needs. On Gemini, choose media resolution per item, so a detailed invoice gets more tokens than a decorative photo. Log tokens per request, so you notice when an innocent change, like full-resolution phone photos, quietly doubles your costs.

Worked example: tenancy applications

A worked example. A property agency in Dubai receives tenancy applications as phone photos, salary letters as PDFs, and screenshots of bank transfers. By policy, passports are excluded. Each application goes into one labelled request, document one the salary letter, image one the transfer screenshot, and the model returns a structured summary with a source label for every field. Anything unreadable or inconsistent goes to an agent. Personal data is minimised before sending, retention is configured to policy, and the provider's data terms were reviewed before launch.

Example 1: flyer vs brand guide

A simple worked example. You want a model to check whether a new flyer follows your two-page brand guide. You send the flyer image and the guide PDF, and ask: does this follow the guide? You get a vague yes. Now restructure the request. Text label: image one is the proposed flyer. The image. The guide as a titled document. Then the question: check image one against the guide for logo placement, colours and fonts; list each issue with the guide page number; if something cannot be judged from the image, say so. The answer now lists two concrete issues, each with a page reference. Same inputs, better packaging.

Example 2: motor insurance claims (illustrative)

Now a business scenario, with illustrative numbers. An insurance claims team in the UK handles about four thousand motor claims a month. Each claim has two to six photos of the damage, a PDF claim form and sometimes a repair estimate. They build one labelled request per claim: form, estimate and photos in a fixed order, each with a label. The model returns a structured summary: vehicle, damage areas seen in photos, whether they match the areas claimed on the form, and the estimate total, with a source label on every field. Mismatches go to a handler. They upload forms through the Files API, downscale photos to what damage assessment needs, and log tokens per claim. Average tokens per claim fall by about a third after downscaling, with no loss in match accuracy on their test set. Illustrative numbers; labelled, ordered, measured.

Pitfalls and metrics

Watch for four pitfalls. Mixing up which image is which because nothing was labelled. Sending full-resolution photos when a cropped region would do. Forgetting request size limits, so production requests fail; handle the error and split inputs. And sending personal or sensitive documents without checking data terms and retention settings. Then measure: field-level accuracy against a labelled set, the rate of correctly flagged unreadable fields, tokens per request, and failures due to size limits.

Recap

To recap. Model APIs take lists of content parts, with different shapes per provider. Put media first, label every item, say what to do with each, and ask for answers that cite labels and pages. Reuse files by ID, cache repeated context, control resolution, and handle errors. Minimise personal data. Try this now: build one mixed-media request for a real task with two images and one PDF, require citations to labels or pages, and log tokens per request. Next: video understanding.

Key takeaways

  • Model APIs accept lists of content parts; provider shapes differ, so learn the image, document and file parts for each.
  • Put media before the question, label every item, and ask for answers that reference item IDs or pages.
  • Use Files APIs, caching and per-item resolution to control cost; handle size-limit errors.
  • Minimise personal data and check provider data terms before sending documents.

Try it

Build one mixed-media request for a real task (two images and one PDF), label every item, require answers that cite item labels or pages, and log tokens per request.