AI · Intermediate · 367 minutes · free · updated Sep 2026
Multimodal & Reasoning Models in Practice
Images, PDFs, video, speech, realtime voice, image and video generation, thinking models and computer-use agents, hands-on
- Lessons
- 17 in 6 modules
- Video lectures
- 17 lectures, 129 minutes
- Updated
- Sep 2026
Tools you'll use
- Claude API
- OpenAI API
- Gemini API
- Veo
- GPT Image
- ElevenLabs
- Python
About this course
AI models now see, listen, speak, generate media, think before answering and operate computers. This intermediate course shows you how to use those capabilities well in real work, with hands-on Python for the Claude, OpenAI and Gemini APIs. You will send images, PDFs and video to models and get reliable, checkable extractions; build transcription and text-to-speech pipelines and realtime voice agents; direct and edit image and video generators with brand consistency, provenance and disclosure in mind; and understand reasoning models, test-time compute and when extended or adaptive thinking is worth the cost. You will learn how computer-use agents work and how to deploy them safely, then choose models by task and benchmark them on your own data. Examples span creators, agencies, SMBs and enterprises in the UK, US, Pakistan, the UAE and Saudi Arabia.
What you will learn
- Send images, PDFs and video to model APIs and get reliable, checkable structured extractions
- Build and evaluate speech-to-text, text-to-speech and realtime voice pipelines
- Direct and edit image and video generators with consistency, provenance and disclosure
- Explain test-time compute and decide when extended or adaptive thinking is worth it
- Tune thinking effort in the Claude, OpenAI and Gemini APIs using evaluation data
- Design computer-use agents with isolation, approvals and injection defences
- Choose models by task and benchmark them on your own data
Course content
Vision: documents, charts, screenshots and UI
Use vision-capable models to read documents, interpret charts and understand screenshots, knowing where they excel and where they fail.
- How vision-language models see — 11 min
- Working with documents, charts, screenshots and UI — 12 min
- Hands-on: sending images and PDFs through model APIs — 13 min
- Video understanding: timestamps, audio and frames — 12 min
Audio: speech-to-text and text-to-speech
Build reliable transcription and voice pipelines, understand their failure modes, and evaluate them properly.
- Speech-to-text: transcription that holds up — 11 min
- Text-to-speech and voice pipelines — 11 min
- Realtime voice agents: speech-to-speech in practice — 13 min
Image and video generation concepts
Understand how image and video generators work conceptually, how to direct them, and the legal and ethical guardrails around them.
- Image generation: how it works and how to direct it — 11 min
- Video generation: concepts, limits and workflows — 11 min
- Image editing, references and brand consistency — 12 min
Reasoning models and test-time compute
Understand what reasoning ('thinking') models do differently, when the extra cost pays off, and how to work with them.
- Reasoning models and test-time compute — 12 min
- Working with reasoning models in practice — 11 min
- Extended and adaptive thinking: controls across APIs — 13 min
Computer-use and browser agents
Understand how agents that see screens and operate browsers or desktops work, where they help, and how to deploy them safely.
Choosing models and benchmarking on your own data
Pick the right model for each task and build lightweight benchmarks on your own data instead of relying on leaderboards.
- Picking models by task — 11 min
- Benchmarking models on your own data — 12 min
Certificate: Certified Multimodal & Reasoning AI Practitioner
The holder can apply vision, speech, image and video generation, reasoning models and computer-use agents to real work. They understand each capability's strengths, failure modes and safeguards, can match reasoning effort to task difficulty, and can select models using criteria and private benchmarks on their own data.
- Final assessment
- 25 questions, 40 minutes
- Passing score
- 80%