Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APIMultimodal inputs and embeddings · Lesson 7 of 19
Multimodal inputs: images, PDFs and audio
Video lecture
Multimodal inputs: images, PDFs and audio
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Multimodal inputs
Most business information doesn't arrive as neat text. It arrives as photos of receipts, PDF rate cards, screenshots of dashboards, and recordings of sales calls. Modern models can read all of these. In this lesson you'll learn how to send images, documents and audio to Claude, OpenAI and Gemini, and how to turn them into reliable data rather than impressive demos.
0:27 Why it matters
Why does this matter? Because the most tedious work in many businesses is reading things and typing them into systems: expenses, invoices, contracts, forms. Multimodal models can take much of that off people's plates. Think of a new assistant who can read handwriting, scan a spreadsheet printout and listen to a voicemail. Useful, but you'd still double check the numbers on their first week. That's the right attitude for multimodal AI too.
0:58 Capabilities (verify per model)
What can each provider handle? All three accept images and PDFs on current models. Claude takes image blocks and document blocks, reading PDF pages as both text and images, with optional citations. OpenAI's Responses API takes input image and input file items, and has dedicated transcription models for audio. Gemini accepts images, PDFs, audio and even video natively inside generate content. Support changes by model, so always check the specific model's documentation.
1:29 Request shapes
Let's look at the request shapes. With Claude, a user message contains a list of blocks: an image block with base64 data and a media type, a document block for the PDF, then a text block with your question. With OpenAI, you send input text and input image items, and you can set the detail level. For audio, OpenAI's transcription endpoint turns a recording into text. With Gemini, you pass parts built from bytes with a MIME type, plus your instruction, in one contents list.
2:06 Simple example: the café receipt
A simple example. You photograph a café receipt and ask a model to extract the merchant, date, total and currency. It works beautifully on a clear photo. Now try a crumpled thermal receipt photographed at an angle in dim light. The model might read a total of eighteen as eighty. The fix isn't hoping for better eyesight. Crop to the receipt, straighten it, and in code check that the line items add up to the total. If they don't, send it to a person.
2:43 Business example: multi-country expenses
Now a realistic business example, with illustrative results. A consultancy with staff in Karachi, Dubai and London receives receipts in rupees, dirhams and pounds, some in Arabic or Urdu. Their pipeline sends each image to a model with a structured output schema: merchant, date, currency, total, tax and line items. Code checks that items add up, the date isn't in the future, and the currency matches the country. Failures go to a finance reviewer. Most receipts pass automatically, and the finance team's time on expenses drops substantially.
3:21 Practical guidance
Some practical guidance. Image quality matters, so crop to what matters and use higher detail settings for small print. Combine vision with structured outputs for any extraction task. Verify numbers in code. Remember that images and PDF pages cost tokens, so a thirty page PDF can be expensive per call: extract only the pages you need, and for documents you reuse, upload them once with a Files API and reference them by id.
3:53 Right tool per document
Here's a cost saving trick many teams miss: not every document needs a multimodal model. A clean digital PDF, one where you can select the text, can be read with an ordinary PDF library, and you send plain text to the model, which is faster and cheaper. Scanned or photographed documents are where multimodal models shine. For long call recordings, transcribe first, then summarize the transcript. Building a fast path for easy documents and a multimodal path for hard ones often cuts cost dramatically without losing accuracy.
4:31 Privacy first
Privacy deserves special attention with multimodal data. Receipts show names and card digits. IDs are highly sensitive. Screenshots reveal whatever was on screen. And call recordings capture other people's voices, which in many places requires notice or consent. Minimize what you send, redact where you can, check each provider's data terms, and involve your privacy team before processing recordings or identity documents at scale.
4:59 Test real scripts
Also test with your real languages and scripts. Accuracy on Arabic or Urdu receipts, handwritten notes or unusual fonts can differ a lot from English printed text. Build your evaluation set from real documents in each language you support, including the bad photos. That's the only way to know how your pipeline performs for all your customers, not just the easy ones.
5:26 Common mistakes
Common mistakes. Sending full resolution photos when a crop would do, which costs more and runs slower. Trusting extracted numbers without arithmetic checks. Uploading call recordings without consent. And assuming audio or video support exists because another model from the same provider has it.
5:45 Video inputs
What about video? Some models, such as Gemini, accept video directly and can describe scenes, read on screen text and answer questions about moments in the footage. For others, a common approach is to sample frames at intervals and transcribe the audio, then send frames plus transcript. For marketing teams, that's useful for checking ad creatives against brand rules or summarizing webinar recordings, but check length limits, costs and data terms first.
6:16 Deeper: 2,000 receipts/month (illustrative)
Let's deepen the multi country expenses example with illustrative numbers. About two thousand receipts a month across Karachi, Dubai and London. In the first month, most were auto approved after passing the sum, date and currency checks. The ones sent to finance were mostly faded thermal receipts and a handful where the model read a date in the wrong format, day and month swapped, which the future date check caught. Finance added one rule: receipts above a threshold always get a human look, regardless of confidence. Staff liked that reimbursement became faster, and finance liked that every automated decision came with the extracted fields and the check results attached.
7:03 Watch me do it: multimodal requests
Watch me do it. Let's walk through the Claude image and PDF request. I read the receipt photo and the rate card PDF from disk and base64 encode each. The user message content is a list of three blocks. First, an image block with a base64 source, the JPEG media type and the data. Second, a document block with the PDF media type. Third, a text block with the question: does the receipt total match the rate card price for the service listed, answer in three bullets. The media blocks come before the text. I call messages create with max tokens fifteen hundred and print the text. The OpenAI version sends input text plus an input image with a data URL and detail high. The transcription example opens the MP3 and calls audio transcriptions create. And the Gemini version passes the audio as a part built from bytes with its MIME type, followed by the summary instruction, in a single contents list.
8:13 Recap + try this now
Quick recap. All three providers read images and PDFs, and audio and video support varies. Use each provider's block or part format, combine vision with structured outputs, check numbers in code, control cost by cropping and reusing uploads, and treat multimodal data as sensitive. Try this now: build a receipt or invoice extractor using structured outputs and a sum check, test it on twenty real documents in the languages you support, and measure field accuracy per language.
What "multimodal" means in practice
Current frontier models accept more than text. Typical capabilities (verify per model):
| Input | Claude | OpenAI | Gemini |
|---|---|---|---|
| Images | Yes (base64, URL or Files API) | Yes (input_image with URL, data URL or file ID) | Yes (inline bytes, uploaded files, URIs) |
| PDFs / documents | Yes (document blocks; pages read as text + images; optional citations) | Yes (input_file) | Yes (PDF as inline data or uploaded file) |
| Audio input | Via separate transcription | Transcription models (e.g., gpt-4o-transcribe) and realtime/audio models | Native audio understanding in generate_content |
| Video input | No (extract frames) | Limited; check docs | Native video understanding |
Use cases: reading receipts and invoices, checking ad creatives against brand guidelines, summarizing call recordings, extracting tables from PDFs, describing product photos for alt text, reviewing screenshots of dashboards.
Request shapes
Claude: image + PDF
import base64, os
import anthropic
client = anthropic.Anthropic()
img_b64 = base64.standard_b64encode(open("receipt.jpg", "rb").read()).decode()
pdf_b64 = base64.standard_b64encode(open("rate-card.pdf", "rb").read()).decode()
msg = client.messages.create(
model=os.environ.get("CLAUDE_MODEL", "claude-sonnet-5"), max_tokens=1500,
messages=[{"role": "user", "content": [
{"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": img_b64}},
{"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": pdf_b64}},
{"type": "text", "text": "Does the receipt total match the rate card price for the service listed? Answer in 3 bullets."},
]}])
print("".join(b.text for b in msg.content if b.type == "text"))For files reused across many requests, upload once with the Files API and reference the file_id instead of resending base64.
OpenAI (Responses API): image
from openai import OpenAI
client = OpenAI()
resp = client.responses.create(
model=os.environ.get("OPENAI_MODEL", "gpt-5.5"),
input=[{"role": "user", "content": [
{"type": "input_text", "text": "List every text claim in this ad creative."},
{"type": "input_image", "image_url": f"data:image/jpeg;base64,{img_b64}", "detail": "high"},
]}])
print(resp.output_text)OpenAI: audio transcription
with open("sales-call.mp3", "rb") as f:
tr = client.audio.transcriptions.create(model="gpt-4o-transcribe", file=f)
print(tr.text)Gemini: image and audio inline
from google import genai
from google.genai import types
g = genai.Client()
resp = g.models.generate_content(
model=os.environ.get("GEMINI_MODEL", "gemini-flash-latest"),
contents=[types.Part.from_bytes(data=open("sales-call.mp3", "rb").read(), mime_type="audio/mp3"),
"Summarize this call: customer needs, objections, next steps."])
print(resp.text)Large files should use each provider's file upload APIs rather than inline bytes; check size limits and supported formats in the docs.
Practical guidance
- Image quality matters: crop to the relevant region, keep text legible, and use higher-detail settings where offered for small print. Rotate photos correctly.
- Ask for structure: combine vision with structured outputs (lesson 6) for extraction tasks.
- Verify numbers: models can misread digits, especially in low-quality scans. Cross-check totals (line items sum to the total) in code.
- Cost: images and PDF pages consume tokens; a 30-page PDF can be expensive per call. Extract only needed pages, and cache repeated documents.
- Privacy: receipts, IDs, call recordings and screenshots often contain personal data. Minimize, redact and check provider data terms and consent (call recordings in particular may require notice to participants under local law).
- Languages: test with your real scripts (Arabic, Urdu); OCR-like accuracy varies by script, font and image quality.
Worked example: expense receipts for a multi-country consultancy
Staff in Karachi, Dubai and London submit photos of receipts in PKR, AED and GBP, some in Arabic or Urdu. Pipeline: image → structured extraction (merchant, date, currency, total, tax, line items) → code checks (line items sum to total within tolerance; date not in the future; currency matches country) → failures to human review. Results (illustrative): most receipts auto-approved; blurred or thermal-faded receipts routed to humans; finance team time on expenses dropped substantially.
Choosing the right tool for the document
Not every document needs a frontier multimodal model. A decision guide:
| Document | Good first choice |
|---|---|
| Clean, digital PDFs with selectable text | Extract text with a PDF library, then send text (cheaper, faster) |
| Scanned or photographed documents | Multimodal model with structured outputs, plus arithmetic checks |
| Complex tables and charts | Multimodal model; ask for tables as structured rows; spot-check against the source |
| Long recordings | Transcribe first, then summarize or extract from the transcript |
| Video | A model with native video support, or sampled frames plus a transcript |
Mixing approaches is normal: a digital-text fast path for most invoices and a multimodal path only for scans can cut cost significantly while keeping accuracy.
Pitfalls
- Sending full-resolution photos when a crop would do (cost and latency).
- Trusting extracted numbers without arithmetic checks.
- Uploading call recordings without consent or notice.
- Assuming video or audio support without checking the specific model.
Measuring success
Field-level accuracy by document type and language, human-review rate, cost per document, and turnaround time.
Key takeaways
- Frontier models accept images and PDFs; audio and video support varies by provider and model.
- Claude uses image and document blocks; OpenAI uses input_image/input_file and transcription models; Gemini accepts inline parts and uploads.
- Combine vision with structured outputs and verify numbers with arithmetic checks in code.
- Images and PDF pages cost tokens: crop, select pages and upload reusable files once.
- Recordings, IDs and receipts contain personal data: minimize, redact and respect consent rules.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Build a receipt or invoice extractor with structured outputs and arithmetic checks, test it on 20 real documents in your languages, and measure field accuracy.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.