Skip to content

AI · Intermediate · 367 minutes · free · updated Sep 2026

Multimodal & Reasoning Models in Practice

Images, PDFs, video, speech, realtime voice, image and video generation, thinking models and computer-use agents, hands-on

Lessons
17 in 6 modules
Video lectures
17 lectures, 129 minutes
Updated
Sep 2026

Tools you'll use

  • Claude API
  • OpenAI API
  • Gemini API
  • Veo
  • GPT Image
  • ElevenLabs
  • Python

Start the course

About this course

AI models now see, listen, speak, generate media, think before answering and operate computers. This intermediate course shows you how to use those capabilities well in real work, with hands-on Python for the Claude, OpenAI and Gemini APIs. You will send images, PDFs and video to models and get reliable, checkable extractions; build transcription and text-to-speech pipelines and realtime voice agents; direct and edit image and video generators with brand consistency, provenance and disclosure in mind; and understand reasoning models, test-time compute and when extended or adaptive thinking is worth the cost. You will learn how computer-use agents work and how to deploy them safely, then choose models by task and benchmark them on your own data. Examples span creators, agencies, SMBs and enterprises in the UK, US, Pakistan, the UAE and Saudi Arabia.

What you will learn

  • Send images, PDFs and video to model APIs and get reliable, checkable structured extractions
  • Build and evaluate speech-to-text, text-to-speech and realtime voice pipelines
  • Direct and edit image and video generators with consistency, provenance and disclosure
  • Explain test-time compute and decide when extended or adaptive thinking is worth it
  • Tune thinking effort in the Claude, OpenAI and Gemini APIs using evaluation data
  • Design computer-use agents with isolation, approvals and injection defences
  • Choose models by task and benchmark them on your own data

Course content

Vision: documents, charts, screenshots and UI

Use vision-capable models to read documents, interpret charts and understand screenshots, knowing where they excel and where they fail.

Audio: speech-to-text and text-to-speech

Build reliable transcription and voice pipelines, understand their failure modes, and evaluate them properly.

Image and video generation concepts

Understand how image and video generators work conceptually, how to direct them, and the legal and ethical guardrails around them.

Reasoning models and test-time compute

Understand what reasoning ('thinking') models do differently, when the extra cost pays off, and how to work with them.

Computer-use and browser agents

Understand how agents that see screens and operate browsers or desktops work, where they help, and how to deploy them safely.

Choosing models and benchmarking on your own data

Pick the right model for each task and build lightweight benchmarks on your own data instead of relying on leaderboards.

Certificate: Certified Multimodal & Reasoning AI Practitioner

The holder can apply vision, speech, image and video generation, reasoning models and computer-use agents to real work. They understand each capability's strengths, failure modes and safeguards, can match reasoning effort to task difficulty, and can select models using criteria and private benchmarks on their own data.

Final assessment
25 questions, 40 minutes
Passing score
80%