Multimodal & Reasoning Models in PracticeVision: documents, charts, screenshots and UI · Lesson 2 of 17
Working with documents, charts, screenshots and UI
Video lecture
Working with documents, charts, screenshots and UI
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Documents, charts, screenshots and UI
Most business uses of vision fall into four jobs: documents, charts, screenshots and user interfaces. Each has its own playbook and its own traps. In this lecture you will learn how to choose a document pipeline, how to send PDFs straight into model APIs with page-level citations, why you should almost never read numbers off a chart image, and how to turn screenshots into triage, QA and even code fixes, while protecting the personal data that images so often contain.
0:35 Document approaches
Start with documents. You have three approaches. Native text extraction for digital PDFs, where the text is already embedded: fast and exact. A vision model on page images for scans, complex layouts, forms and tables. Or a hybrid: extract text where you can, and send page images for pages with tables, stamps or diagrams. Most major APIs now accept PDFs directly and process both the text layer and page images, which is convenient. For long documents, process page ranges and combine, and keep page numbers in every output for traceability.
1:14 PDFs through the API
Here is the hands-on version with Claude. You send the PDF as a document block, from base64, a URL, or a Files API ID if you reuse it. You give it a title and enable citations. Then you ask your question, for example the notice period and any late-payment penalties, answering not stated if the contract does not say. The response comes back as text blocks, and cited blocks carry page locations, so you can print the answer with page numbers. OpenAI's Responses API takes an input file part, and Gemini takes the PDF as a file part with media resolution controlling tokens per page. Request size and page limits apply, so split very long documents.
2:04 Charts: the golden rule
Now charts. Vision models are good at describing trends, comparing series and summarising what a chart shows. They are much less reliable at exact values, especially with dense axes, log scales, dual axes or overlapping series. So here is the single most important rule: if you have the underlying data, use the data, not the chart image. When you only have the image, ask for qualitative readings, like which region grew fastest, and if you need numbers, ask for approximate values with stated uncertainty.
2:41 Charts: a misleading-design check
Vision models also make useful chart reviewers. Ask one to describe a chart for a non-technical manager in three bullets, and then to list any features that could mislead a reader: a truncated y-axis, dual axes, missing labels, a stacked chart that hides a decline. Run that check before any board deck or client report goes out. It catches the chart choices that quietly erode trust.
3:10 Screenshots in practice
Screenshots are some of the most practical inputs. In support, a customer's error screenshot can be classified, the error text transcribed, and that text used to retrieve the right help article. In QA, you can compare a screenshot to the design spec and list differences. In documentation, a sequence of screenshots becomes step-by-step instructions. Tips: crop out irrelevant areas, redact personal data such as names, emails and account numbers before sending, and tell the model the product and version.
3:44 UI review and code
User interfaces go one step further. Models can review a screen against explicit criteria: is the primary action obvious, are errors and required fields clear, are there visible accessibility issues like low contrast or tiny tap targets. They can also turn a screenshot or sketch into working front-end code, and propose code changes for the top issues. Treat that as a draft for a developer. And note the limit: a static image cannot reveal keyboard navigation, focus order, screen-reader labels, or exact contrast ratios. So ask the model to mark those as needs testing, and verify with proper accessibility tools.
4:27 Privacy and automation decisions
Images often carry more personal data than text: faces, names on screens, addresses on parcels, ID documents. Before building a vision pipeline, minimise: crop or redact what you do not need. Check your provider's data retention and training terms, and your contracts. Be especially careful with identity documents, biometric data and images of people, which are sensitive under many data protection laws, including the UK GDPR, the UAE's data protection law and Saudi Arabia's personal data protection law. And decide automation field by field: totals and dates may be reliable enough to automate, while handwritten notes should always go to review.
5:11 Example 1: a Q3 sales chart
A simple worked example with a chart. Your manager forwards a screenshot of a bar chart and asks: how much did sales grow in Q3? Ask a vision model directly, and it reads roughly forty percent off the axis. But look closer: the y-axis starts at eighty, not zero, so the bars exaggerate the change. The better move: ask the model to describe the trend in plain words and to list any misleading features, then get the real numbers from the spreadsheet behind the chart. The model catches the truncated axis. The spreadsheet shows growth of twelve percent. Use the model for reading the design, and the data for the numbers.
5:59 Example 2: contract extraction (illustrative)
Now a business scenario, with illustrative numbers. An accounting firm in Riyadh receives about eight hundred supplier contracts a year as PDFs, some scanned, some digital, in Arabic and English. They need five fields from each: parties, start date, end date, renewal terms and termination notice. They test direct PDF input with citations on sixty contracts. Digital PDFs score about ninety-seven percent field accuracy, scanned ones about eighty-five percent, with most errors in dates on stamped pages. So they route by type: digital contracts go straight through with citations and a spot-check of one in ten; scanned contracts get page images at higher resolution, and any date field on a stamped page goes to review. Every extracted field carries its page number, so reviewers jump straight to the evidence. Illustrative numbers, but automating field by field and document type by type is the durable lesson.
7:02 Recap
To recap. Choose native text, vision or hybrid for documents, keep page numbers, and use direct PDF input with citations where available. For charts, use the underlying data whenever you can, and use models to catch misleading design. Screenshots power triage, QA and documentation, and UI review complements, but never replaces, accessibility testing. Minimise personal data. Try this now: take one chart from a recent report, ask a vision model to describe it and list misleading features, and compare its reading against the underlying data. Next: hands-on multimodal API requests.
Four common jobs, four playbooks
Most business uses of vision fall into four jobs. Each has its own techniques and failure modes.
1. Documents: PDFs, scans and forms
Approaches:
- Native text extraction for digital PDFs (text is already embedded). Fast and exact; use it when available.
- Vision model on page images for scans, complex layouts, or when layout matters (forms, tables).
- Hybrid: extract text where possible and send page images for pages with tables, stamps or diagrams.
Many model APIs accept PDFs directly and process both text and page images; check how your provider handles them and what it costs per page.
Techniques:
- Process page by page for long documents, then combine (map-reduce).
- For tables, ask for output in a structured format with the header row preserved, then validate totals in code.
- Keep page numbers in outputs for traceability.
2. Charts and data visualisations
Vision models can describe trends, compare series and summarise what a chart shows. They are less reliable at exact values, especially with dense axes, log scales or overlapping series.
Playbook:
- If you have the underlying data, use the data, not the chart image. This is the single most important rule.
- Ask for qualitative readings ("Which region grew fastest?") rather than precise numbers.
- When numbers are needed, ask the model to state approximate values with explicit uncertainty ("about 40, reading from the axis").
- Watch for misreadings: dual axes, truncated y-axes, stacked charts.
Describe what this chart shows for a non-technical manager in 3 bullets.
Then list any features that could mislead a reader (e.g. truncated axis,
dual axes, missing labels). Do not estimate exact values.This also makes vision models useful as chart reviewers, catching misleading visual choices before a presentation.
3. Screenshots
Screenshots of apps, dashboards, error messages and websites are some of the most practical inputs.
Uses:
- Support: a customer sends a screenshot of an error; the model identifies the screen and likely cause from your help articles.
- QA: compare a screenshot to the design specification and list differences.
- Documentation: generate step-by-step instructions from a sequence of screenshots.
- Competitive analysis: summarise features visible on a competitor's public pages.
Tips: crop out irrelevant areas, redact personal data before sending (names, emails, account numbers), and state the product and version.
4. User interfaces
Beyond reading a screen, models can reason about UI: accessibility issues (low contrast, missing labels), usability heuristics, and consistency with a design system. They can also generate front-end code from a screenshot or sketch; results are useful drafts that need developer review.
Review this checkout screen against these criteria:
1. Is the primary action obvious?
2. Are error states and required fields clear?
3. Any likely accessibility issues visible (contrast, tiny tap targets,
text in images)?
For each issue, describe its location on the screen and a suggested fix.
Mark anything you cannot judge from a static image as "needs testing".Notice the last line: a static image cannot reveal keyboard navigation, screen-reader labels or real contrast values precisely. Vision review complements, not replaces, proper accessibility testing.
Worked example: a support triage flow
A software company receives many support emails with screenshots. Their pipeline:
- A vision model classifies the screenshot (which screen, error visible or not, error text transcribed).
- The transcribed error text drives retrieval over help articles (RAG).
- A draft reply cites the relevant article.
- Screenshots containing personal data are flagged and stored according to policy.
Agents resolve more tickets on first reply, and the error text extraction also produces a useful dataset of the most common errors for the product team.
Privacy and compliance
Images often contain more personal data than text: faces, names on screens, addresses on parcels, ID documents. Before building a vision pipeline:
- Minimise: crop or redact what is not needed.
- Check your provider's data retention and training policies and your contractual terms.
- Be careful with identity documents, biometric data and images of people, which are sensitive under many data protection laws. Seek appropriate advice.
Hands-on: PDFs straight into the API
Most major APIs now accept PDFs directly and process both the text layer and page images, which helps with tables, charts and scans.
Claude: send the PDF as a document block (base64, a URL, or a Files API ID for reuse across requests). Enable citations to get answers tied to page numbers.
import base64, os, anthropic
client = anthropic.Anthropic()
pdf_b64 = base64.standard_b64encode(open("contracts/supplier-2026.pdf", "rb").read()).decode()
resp = client.messages.create(
model=os.environ.get("CLAUDE_MODEL", "claude-opus-5"), max_tokens=2048,
messages=[{"role": "user", "content": [
{"type": "document",
"source": {"type": "base64", "media_type": "application/pdf", "data": pdf_b64},
"title": "Supplier contract 2026", "citations": {"enabled": True}},
{"type": "text", "text": "What is the notice period to terminate, and are there late-payment "
"penalties? If the contract does not say, answer 'Not stated'."},
]}],
)
for block in resp.content:
if block.type == "text":
pages = [f"p.{c.start_page_number}" for c in (block.citations or [])
if getattr(c, "start_page_number", None)]
print(block.text, pages)OpenAI (Responses API): include an input_file content item (uploaded file ID or inline data) alongside input_text. Gemini: pass the PDF as a file part; media_resolution controls tokens per page.
Request size and page limits apply (and differ by model and provider), so split very long documents and process them page range by page range.
Choosing a document pipeline
| Document type | Recommended approach |
|---|---|
| Digital PDF, mostly text | Native text extraction, or direct PDF input for convenience |
| Scans and photos | Vision model on page images; route low-readability pages to review |
| Tables with totals | Structured output plus code checks that rows sum correctly |
| Very long reports | Page-range map-reduce with page numbers kept throughout |
| Forms at high volume | Compare a vision model against OCR plus a text model on cost and accuracy |
Screenshots and UI: generate fixes, not just critiques
Vision models can now turn a screenshot or sketch into working front-end code and can compare a build against a design. A productive workflow for small teams: screenshot the page, ask for issues against explicit criteria (contrast, hierarchy, tap-target size, error states), then ask for code changes for the top three. Treat the output as a draft for a developer, and verify accessibility with proper tools (automated checkers and screen-reader testing), because a static image cannot reveal focus order or real contrast ratios precisely.
Going further
For high-volume document pipelines, build a golden set per document type (invoice, contract, ID form) with field-level labels, and measure accuracy per field. Some fields (totals, dates) may be reliable enough to automate, while others (handwritten notes) should always go to review. Automation decisions should be made field by field, not document by document.
Key takeaways
- Use native text for digital PDFs, vision for scans and layouts, or a hybrid; keep page numbers.
- For charts, use the underlying data when available; ask vision models for qualitative readings and misleading-design checks.
- Screenshots support triage, QA and documentation; crop and redact personal data first.
- UI review from images complements, but does not replace, real accessibility and usability testing.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Take one chart from a recent report. Ask a vision model to describe it and list misleading features. Compare its reading against the underlying data.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.