Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APIMultimodal inputs and embeddings · Lesson 7 of 19

Multimodal inputs: images, PDFs and audio

Article · 16 min · 9 min lecture

Video lecture

Multimodal inputs: images, PDFs and audio

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Multimodal inputs

  • Images, PDFs, audio
  • Provider request shapes
  • From demo to reliable data

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What "multimodal" means in practice

Current frontier models accept more than text. Typical capabilities (verify per model):

InputClaudeOpenAIGemini
ImagesYes (base64, URL or Files API)Yes (input_image with URL, data URL or file ID)Yes (inline bytes, uploaded files, URIs)
PDFs / documentsYes (document blocks; pages read as text + images; optional citations)Yes (input_file)Yes (PDF as inline data or uploaded file)
Audio inputVia separate transcriptionTranscription models (e.g., gpt-4o-transcribe) and realtime/audio modelsNative audio understanding in generate_content
Video inputNo (extract frames)Limited; check docsNative video understanding

Use cases: reading receipts and invoices, checking ad creatives against brand guidelines, summarizing call recordings, extracting tables from PDFs, describing product photos for alt text, reviewing screenshots of dashboards.

Request shapes

Claude: image + PDF

import base64, os
import anthropic

client = anthropic.Anthropic()
img_b64 = base64.standard_b64encode(open("receipt.jpg", "rb").read()).decode()
pdf_b64 = base64.standard_b64encode(open("rate-card.pdf", "rb").read()).decode()

msg = client.messages.create(
    model=os.environ.get("CLAUDE_MODEL", "claude-sonnet-5"), max_tokens=1500,
    messages=[{"role": "user", "content": [
        {"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": img_b64}},
        {"type": "document", "source": {"type": "base64", "media_type": "application/pdf", "data": pdf_b64}},
        {"type": "text", "text": "Does the receipt total match the rate card price for the service listed? Answer in 3 bullets."},
    ]}])
print("".join(b.text for b in msg.content if b.type == "text"))

For files reused across many requests, upload once with the Files API and reference the file_id instead of resending base64.

OpenAI (Responses API): image

from openai import OpenAI
client = OpenAI()
resp = client.responses.create(
    model=os.environ.get("OPENAI_MODEL", "gpt-5.5"),
    input=[{"role": "user", "content": [
        {"type": "input_text", "text": "List every text claim in this ad creative."},
        {"type": "input_image", "image_url": f"data:image/jpeg;base64,{img_b64}", "detail": "high"},
    ]}])
print(resp.output_text)

OpenAI: audio transcription

with open("sales-call.mp3", "rb") as f:
    tr = client.audio.transcriptions.create(model="gpt-4o-transcribe", file=f)
print(tr.text)

Gemini: image and audio inline

from google import genai
from google.genai import types
g = genai.Client()
resp = g.models.generate_content(
    model=os.environ.get("GEMINI_MODEL", "gemini-flash-latest"),
    contents=[types.Part.from_bytes(data=open("sales-call.mp3", "rb").read(), mime_type="audio/mp3"),
              "Summarize this call: customer needs, objections, next steps."])
print(resp.text)

Large files should use each provider's file upload APIs rather than inline bytes; check size limits and supported formats in the docs.

Practical guidance

  • Image quality matters: crop to the relevant region, keep text legible, and use higher-detail settings where offered for small print. Rotate photos correctly.
  • Ask for structure: combine vision with structured outputs (lesson 6) for extraction tasks.
  • Verify numbers: models can misread digits, especially in low-quality scans. Cross-check totals (line items sum to the total) in code.
  • Cost: images and PDF pages consume tokens; a 30-page PDF can be expensive per call. Extract only needed pages, and cache repeated documents.
  • Privacy: receipts, IDs, call recordings and screenshots often contain personal data. Minimize, redact and check provider data terms and consent (call recordings in particular may require notice to participants under local law).
  • Languages: test with your real scripts (Arabic, Urdu); OCR-like accuracy varies by script, font and image quality.

Worked example: expense receipts for a multi-country consultancy

Staff in Karachi, Dubai and London submit photos of receipts in PKR, AED and GBP, some in Arabic or Urdu. Pipeline: image → structured extraction (merchant, date, currency, total, tax, line items) → code checks (line items sum to total within tolerance; date not in the future; currency matches country) → failures to human review. Results (illustrative): most receipts auto-approved; blurred or thermal-faded receipts routed to humans; finance team time on expenses dropped substantially.

Choosing the right tool for the document

Not every document needs a frontier multimodal model. A decision guide:

DocumentGood first choice
Clean, digital PDFs with selectable textExtract text with a PDF library, then send text (cheaper, faster)
Scanned or photographed documentsMultimodal model with structured outputs, plus arithmetic checks
Complex tables and chartsMultimodal model; ask for tables as structured rows; spot-check against the source
Long recordingsTranscribe first, then summarize or extract from the transcript
VideoA model with native video support, or sampled frames plus a transcript

Mixing approaches is normal: a digital-text fast path for most invoices and a multimodal path only for scans can cut cost significantly while keeping accuracy.

Pitfalls

  • Sending full-resolution photos when a crop would do (cost and latency).
  • Trusting extracted numbers without arithmetic checks.
  • Uploading call recordings without consent or notice.
  • Assuming video or audio support without checking the specific model.

Measuring success

Field-level accuracy by document type and language, human-review rate, cost per document, and turnaround time.

Key takeaways

  • Frontier models accept images and PDFs; audio and video support varies by provider and model.
  • Claude uses image and document blocks; OpenAI uses input_image/input_file and transcription models; Gemini accepts inline parts and uploads.
  • Combine vision with structured outputs and verify numbers with arithmetic checks in code.
  • Images and PDF pages cost tokens: crop, select pages and upload reusable files once.
  • Recordings, IDs and receipts contain personal data: minimize, redact and respect consent rules.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An expense extractor sometimes misreads totals on faded receipts. What is the best safeguard?
  2. You will analyze the same 40-page policy PDF across hundreds of requests. What should you do?
  3. Before transcribing sales calls, what must you check?

Put it into practice

Build a receipt or invoice extractor with structured outputs and arithmetic checks, test it on 20 real documents in your languages, and measure field accuracy.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.