Multimodal & Reasoning Models in PracticeChoosing models and benchmarking on your own data · Lesson 16 of 17
Picking models by task
Video lecture
Picking models by task
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Picking models by task
Which model is the best? It is the most common question in AI, and the wrong one. The best model for Urdu social captions may not be the best for invoice extraction, and neither may be the best for a real-time voice agent. In this lecture you will learn the criteria that matter, how to shortlist in a fast-moving market, proprietary versus open-weight trade-offs, how to match tasks to model types, how to list models programmatically, and the portfolio approach most teams end up using.
0:37 Why model choice matters
Why does this matter? Because model choice quietly shapes quality, cost and risk for every AI feature you run. Choose by brand or leaderboard, and you may overpay for simple tasks, underperform on the languages your customers speak, or discover too late that the data terms do not fit. Think of it like hiring. You would not hire the same person for accounting, design and customer calls just because they topped one exam.
1:09 Ten criteria
There is no single best model. The market offers large and small, general and specialised, reasoning and fast, text-only and multimodal, proprietary and open-weight. Rankings shift with every release. The durable skill is knowing which properties matter for your task and testing candidates against them. Ten criteria cover most decisions: capability fit, modalities, context length, latency, cost, reliability features like structured outputs and tool calling, quality in your users' languages, data and compliance terms, deployment options, and vendor factors like deprecation policies and rate limits.
1:46 Shortlisting
In twenty twenty-six, the major closed-model families include Anthropic's Claude, with Opus, Sonnet and Haiku tiers; OpenAI's GPT models; and Google's Gemini, with Pro, Flash and Flash-Lite tiers; alongside strong open-weight families from several labs. Names and versions change every few months. So shortlist in three steps. Apply hard filters first: required inputs and outputs, context length, data residency and contract terms, and platform availability. Then tier by task: a flagship for hard reasoning, a mid-tier model for most work, a fast tier for high-volume simple steps. Finally, put two or three candidates per task into your private benchmark.
2:29 Proprietary vs open-weight
Proprietary hosted models often lead on frontier capability and come with managed infrastructure, safety systems and features. Open-weight models, with downloadable weights under a licence, offer control, self-hosting for residency or privacy, and fine-tuning flexibility, but require infrastructure, security and operational expertise. Check licences carefully, because open licences vary in commercial terms. Many organisations use both: hosted frontier models for hard tasks, and smaller or self-hosted models for high-volume or sensitive workloads.
3:00 Task → priorities
Match tasks to priorities. Customer chat: latency, tone, tool calling, safety and languages. Document extraction: vision and PDF handling, structured outputs, and cost per page. Research and analysis: reasoning, long context and citations. Coding: reasoning, tool use and code quality on your stack. Bulk classification: cost, throughput and batch support; small models often suffice. Voice agents: streaming speech quality and end-to-end latency. Image and video creation: prompt adherence, style control and licensing terms.
3:32 Stay current
Keep your view current with code. The lesson's snippet lists models from the Anthropic and OpenAI APIs, and the Anthropic listing includes context window information. Run it in a scheduled job that alerts you when new models appear or old ones disappear, and cross-check the providers' deprecation pages. Then fill in a capability matrix for your shortlist, image input, PDF input with citations, audio and video input, image output, structured outputs, effort controls, computer use, and data residency, from official documentation, with the date you checked each cell.
4:10 The portfolio
Most teams end up with a portfolio: a default capable model for most tasks, a fast, cheap model for high-volume simple steps, a reasoning configuration for hard cases, and specialised models for embeddings, speech or image generation. Put an abstraction layer in your code, a single internal interface that can call different providers, and keep prompts and evaluation sets per task, so switching is a measured decision. A marketing agency serving Pakistan, the UAE and the UK tested three models on thirty real briefs per language. One clearly led on Urdu, another on Arabic, so they route by language.
4:53 Example 1: a solo creator's three tasks
A simple worked example. A solo creator uses AI for three things: writing video scripts, turning long recordings into captions, and generating thumbnail ideas. One model for everything seems simplest. But the needs differ. Scripts need good writing and a consistent voice. Captions need accurate speech-to-text with timestamps, which is a specialised model. Thumbnails need an image model with good text rendering and a commercial-use licence. Picking by task, three tools instead of one, gives better results and often costs less, because the captioning model is far cheaper than a flagship chat model.
5:33 Example 2: a healthcare group (illustrative)
Now a business scenario, with illustrative numbers. A healthcare group in the UAE wants AI for three workflows: summarising patient-feedback surveys, extracting data from referral letters, and an internal policy assistant for staff. Hard filters come first. Patient-related data must be processed under their data residency and contractual requirements, which rules out some deployment options. For referral extraction they shortlist two vision-capable models available on an approved cloud region with structured outputs. For survey summaries, a fast, low-cost tier is enough. For the policy assistant, they need long context and citations. They benchmark two candidates per workflow on their own data and record everything in a model register with pinned versions and deprecation dates. Six months later a candidate is retired; the switch takes days, not months. Illustrative scenario.
6:29 Traps and the register
Avoid four traps. Choosing by leaderboard rank or brand alone. Ignoring total cost, including retries, reasoning tokens and human review. Forgetting data terms until procurement blocks deployment. And locking into one provider's proprietary features without an exit plan. Keep a simple model register: which models serve which tasks, pinned versions, evaluation scores, costs, data terms and deprecation dates. Review it quarterly. That is how you keep pace without constant disruption.
6:59 Recap
To recap. There is no single best model; choose by the criteria that matter for each task. Shortlist with hard filters, tiers and a few candidates, weigh proprietary against open-weight, and keep your view current with model listings and a dated capability matrix. Run a portfolio behind an abstraction layer, and keep a model register. Try this now: list three AI tasks in your organisation, write the top three selection criteria for each, and name one candidate model type to test. Next: benchmarking models on your own data.
No single best model
The market offers many models: large and small, general and specialised, reasoning and fast, text-only and multimodal, proprietary and open-weight. Rankings shift frequently with new releases. The durable skill is not memorising which model leads this month, but knowing which properties matter for your task and testing candidates against them.
The selection criteria
- Capability fit: can it do the task well? (Test, don't assume.)
- Modalities: does it accept the inputs you need (images, PDFs, audio, video) and produce the outputs you need?
- Context length: can it take your typical input size, with quality holding up at that length?
- Latency: time to first token and total time, including reasoning.
- Cost: per-token pricing, reasoning tokens, image or audio pricing, caching and batch discounts.
- Reliability features: structured outputs, tool calling quality, citation support.
- Languages: quality in your users' languages, not just English.
- Data and compliance: data retention, training use, regional hosting or data residency options, certifications, contractual terms.
- Deployment options: hosted API, cloud marketplace, or self-hosted open-weight models.
- Vendor factors: stability, deprecation policies, support and rate limits.
Proprietary vs open-weight
Proprietary hosted models often lead on frontier capability and come with managed infrastructure, safety systems and features. Open-weight models (weights downloadable under a licence) offer control, the ability to self-host for data residency or privacy, and fine-tuning flexibility, but require infrastructure, security and operational expertise. Check licences carefully; "open" licences vary in commercial terms.
Many organisations use both: hosted frontier models for hard tasks, smaller or self-hosted models for high-volume or sensitive workloads.
Matching tasks to model types
| Task | What to prioritise |
|---|---|
| Customer chat | Latency, tone, tool calling, safety, languages |
| Document extraction | Vision and PDF handling, structured outputs, cost per page |
| Research and analysis | Reasoning, long context, citations |
| Coding assistance | Reasoning, tool use, code quality on your stack |
| Bulk classification | Cost, throughput, batch support; small models often suffice |
| Voice agents | Streaming STT/TTS quality, end-to-end latency |
| Image/video creation | Prompt adherence, style control, licensing terms |
The portfolio approach
Rather than one model for everything, many teams run a portfolio:
- A default capable model for most tasks.
- A fast, cheap model for high-volume simple steps.
- A reasoning configuration for hard cases.
- Specialised models where needed (embeddings, STT, image generation).
An abstraction layer in your code (a single internal interface that can call different providers) makes switching easier as the market changes. Keep prompts and evaluation sets per task so switching is a measured decision.
Worked example
A marketing agency serving clients in Pakistan, the UAE and the UK chooses models for four workflows:
- Social caption drafting in English, Urdu and Arabic: they test three models on 30 real briefs per language; one clearly leads on Urdu, another on Arabic, so they route by language.
- Invoice extraction: a vision-capable model with structured outputs; cost per page is the deciding factor between two similar performers.
- Campaign analysis: a reasoning model for quarterly deep dives only.
- Comment moderation: a small, fast model with escalation of uncertain cases.
Traps
- Choosing by leaderboard rank or brand alone.
- Ignoring total cost (retries, reasoning tokens, human review).
- Forgetting data terms until procurement blocks deployment.
- Locking into one provider's proprietary features without an exit plan.
How to shortlist in a fast-moving market
The major closed-model families in 2026 are Anthropic's Claude (Opus, Sonnet and Haiku tiers), OpenAI's GPT models, and Google's Gemini (Pro, Flash and Flash-Lite tiers), alongside strong open-weight families from several labs. Names and versions change every few months, so build your shortlist from the providers' current model pages and the Models APIs, which report context windows and capabilities programmatically.
A practical shortlisting sequence:
- Hard filters first: required modalities (PDF, audio, video in; images or speech out), context length, data residency and contractual terms, platform availability (direct API, AWS, Google Cloud, Azure).
- Tier by task: a flagship model for hard reasoning, a mid-tier model for the bulk of work, a fast tier for high-volume simple steps.
- Two or three candidates per task into your private benchmark (next lesson).
Hands-on: list models and capabilities programmatically
import anthropic
from openai import OpenAI
for m in anthropic.Anthropic().models.list():
print("anthropic", m.id, getattr(m, "max_input_tokens", None))
for m in OpenAI().models.list():
print("openai", m.id)Use this in a scheduled job that alerts you when new models appear or old ones disappear, and cross-check deprecation pages. The Gemini API offers a similar models listing in its SDK.
Multimodal capability matrix (fill in for your shortlist)
Capability | Candidate A | Candidate B | Candidate C
----------------------|-------------|-------------|------------
Image input | | |
PDF input + citations | | |
Audio input | | |
Video input | | |
Image output | | |
Structured outputs | | |
Effort / thinking ctl | | |
Computer use | | |
Data residency | | |Fill every cell from official documentation, with the date you checked. That table, plus your benchmark scores, is a decision record you can defend.
Going further
Maintain a simple model register: which models are used for which tasks, versions pinned, evaluation scores, costs, data terms and deprecation dates. Review it quarterly; this is how you keep pace with a fast-moving market without constant disruption.
Key takeaways
- There is no single best model; choose by capability fit, modalities, context, latency, cost, reliability features, languages, compliance and deployment.
- Proprietary hosted models and open-weight models each have trade-offs; many teams use both.
- Run a portfolio: default, fast/cheap, reasoning and specialised models, behind an abstraction layer.
- Avoid leaderboard-only decisions; keep a model register with scores, costs, data terms and deprecations.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
List three AI tasks in your organisation. For each, write the top three selection criteria from this lesson and one candidate model type to test.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.