Multimodal & Reasoning Models in PracticeVision: documents, charts, screenshots and UI · Lesson 2 of 17

Working with documents, charts, screenshots and UI

Article · 12 min · 8 min lecture

Video lecture

Working with documents, charts, screenshots and UI

11 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 11

Documents, charts, screenshots and UI

  • Four jobs, four playbooks
  • PDFs with page citations
  • Charts: use the data
  • Screenshots and UI

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Four common jobs, four playbooks

Most business uses of vision fall into four jobs. Each has its own techniques and failure modes.

1. Documents: PDFs, scans and forms

Approaches:

  • Native text extraction for digital PDFs (text is already embedded). Fast and exact; use it when available.
  • Vision model on page images for scans, complex layouts, or when layout matters (forms, tables).
  • Hybrid: extract text where possible and send page images for pages with tables, stamps or diagrams.

Many model APIs accept PDFs directly and process both text and page images; check how your provider handles them and what it costs per page.

Techniques:

  • Process page by page for long documents, then combine (map-reduce).
  • For tables, ask for output in a structured format with the header row preserved, then validate totals in code.
  • Keep page numbers in outputs for traceability.

2. Charts and data visualisations

Vision models can describe trends, compare series and summarise what a chart shows. They are less reliable at exact values, especially with dense axes, log scales or overlapping series.

Playbook:

  • If you have the underlying data, use the data, not the chart image. This is the single most important rule.
  • Ask for qualitative readings ("Which region grew fastest?") rather than precise numbers.
  • When numbers are needed, ask the model to state approximate values with explicit uncertainty ("about 40, reading from the axis").
  • Watch for misreadings: dual axes, truncated y-axes, stacked charts.
Describe what this chart shows for a non-technical manager in 3 bullets.
Then list any features that could mislead a reader (e.g. truncated axis,
dual axes, missing labels). Do not estimate exact values.

This also makes vision models useful as chart reviewers, catching misleading visual choices before a presentation.

3. Screenshots

Screenshots of apps, dashboards, error messages and websites are some of the most practical inputs.

Uses:

  • Support: a customer sends a screenshot of an error; the model identifies the screen and likely cause from your help articles.
  • QA: compare a screenshot to the design specification and list differences.
  • Documentation: generate step-by-step instructions from a sequence of screenshots.
  • Competitive analysis: summarise features visible on a competitor's public pages.

Tips: crop out irrelevant areas, redact personal data before sending (names, emails, account numbers), and state the product and version.

4. User interfaces

Beyond reading a screen, models can reason about UI: accessibility issues (low contrast, missing labels), usability heuristics, and consistency with a design system. They can also generate front-end code from a screenshot or sketch; results are useful drafts that need developer review.

Review this checkout screen against these criteria:
1. Is the primary action obvious?
2. Are error states and required fields clear?
3. Any likely accessibility issues visible (contrast, tiny tap targets,
   text in images)?
For each issue, describe its location on the screen and a suggested fix.
Mark anything you cannot judge from a static image as "needs testing".

Notice the last line: a static image cannot reveal keyboard navigation, screen-reader labels or real contrast values precisely. Vision review complements, not replaces, proper accessibility testing.

Worked example: a support triage flow

A software company receives many support emails with screenshots. Their pipeline:

  1. A vision model classifies the screenshot (which screen, error visible or not, error text transcribed).
  2. The transcribed error text drives retrieval over help articles (RAG).
  3. A draft reply cites the relevant article.
  4. Screenshots containing personal data are flagged and stored according to policy.

Agents resolve more tickets on first reply, and the error text extraction also produces a useful dataset of the most common errors for the product team.

Privacy and compliance

Images often contain more personal data than text: faces, names on screens, addresses on parcels, ID documents. Before building a vision pipeline:

  • Minimise: crop or redact what is not needed.
  • Check your provider's data retention and training policies and your contractual terms.
  • Be careful with identity documents, biometric data and images of people, which are sensitive under many data protection laws. Seek appropriate advice.

Hands-on: PDFs straight into the API

Most major APIs now accept PDFs directly and process both the text layer and page images, which helps with tables, charts and scans.

Claude: send the PDF as a document block (base64, a URL, or a Files API ID for reuse across requests). Enable citations to get answers tied to page numbers.

import base64, os, anthropic
client = anthropic.Anthropic()

pdf_b64 = base64.standard_b64encode(open("contracts/supplier-2026.pdf", "rb").read()).decode()
resp = client.messages.create(
    model=os.environ.get("CLAUDE_MODEL", "claude-opus-5"), max_tokens=2048,
    messages=[{"role": "user", "content": [
        {"type": "document",
         "source": {"type": "base64", "media_type": "application/pdf", "data": pdf_b64},
         "title": "Supplier contract 2026", "citations": {"enabled": True}},
        {"type": "text", "text": "What is the notice period to terminate, and are there late-payment "
         "penalties? If the contract does not say, answer 'Not stated'."},
    ]}],
)
for block in resp.content:
    if block.type == "text":
        pages = [f"p.{c.start_page_number}" for c in (block.citations or [])
                 if getattr(c, "start_page_number", None)]
        print(block.text, pages)

OpenAI (Responses API): include an input_file content item (uploaded file ID or inline data) alongside input_text. Gemini: pass the PDF as a file part; media_resolution controls tokens per page.

Request size and page limits apply (and differ by model and provider), so split very long documents and process them page range by page range.

Choosing a document pipeline

Document typeRecommended approach
Digital PDF, mostly textNative text extraction, or direct PDF input for convenience
Scans and photosVision model on page images; route low-readability pages to review
Tables with totalsStructured output plus code checks that rows sum correctly
Very long reportsPage-range map-reduce with page numbers kept throughout
Forms at high volumeCompare a vision model against OCR plus a text model on cost and accuracy

Screenshots and UI: generate fixes, not just critiques

Vision models can now turn a screenshot or sketch into working front-end code and can compare a build against a design. A productive workflow for small teams: screenshot the page, ask for issues against explicit criteria (contrast, hierarchy, tap-target size, error states), then ask for code changes for the top three. Treat the output as a draft for a developer, and verify accessibility with proper tools (automated checkers and screen-reader testing), because a static image cannot reveal focus order or real contrast ratios precisely.

Going further

For high-volume document pipelines, build a golden set per document type (invoice, contract, ID form) with field-level labels, and measure accuracy per field. Some fields (totals, dates) may be reliable enough to automate, while others (handwritten notes) should always go to review. Automation decisions should be made field by field, not document by document.

Key takeaways

  • Use native text for digital PDFs, vision for scans and layouts, or a hybrid; keep page numbers.
  • For charts, use the underlying data when available; ask vision models for qualitative readings and misleading-design checks.
  • Screenshots support triage, QA and documentation; crop and redact personal data first.
  • UI review from images complements, but does not replace, real accessibility and usability testing.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. You have both a chart image and the spreadsheet behind it. What should the AI analyse for exact figures?
  2. Which UI issue can a vision model NOT reliably judge from a static screenshot?
  3. What should you do before sending customer screenshots to an external AI service?

Put it into practice

Take one chart from a recent report. Ask a vision model to describe it and list misleading features. Compare its reading against the underlying data.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.