Multimodal & Reasoning Models in PracticeChoosing models and benchmarking on your own data · Lesson 16 of 17

Picking models by task

Article · 11 min · 8 min lecture

Video lecture

Picking models by task

12 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 12

Picking models by task

  • Selection criteria
  • Shortlisting in a fast market
  • Proprietary vs open-weight
  • The portfolio approach

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

No single best model

The market offers many models: large and small, general and specialised, reasoning and fast, text-only and multimodal, proprietary and open-weight. Rankings shift frequently with new releases. The durable skill is not memorising which model leads this month, but knowing which properties matter for your task and testing candidates against them.

The selection criteria

  1. Capability fit: can it do the task well? (Test, don't assume.)
  2. Modalities: does it accept the inputs you need (images, PDFs, audio, video) and produce the outputs you need?
  3. Context length: can it take your typical input size, with quality holding up at that length?
  4. Latency: time to first token and total time, including reasoning.
  5. Cost: per-token pricing, reasoning tokens, image or audio pricing, caching and batch discounts.
  6. Reliability features: structured outputs, tool calling quality, citation support.
  7. Languages: quality in your users' languages, not just English.
  8. Data and compliance: data retention, training use, regional hosting or data residency options, certifications, contractual terms.
  9. Deployment options: hosted API, cloud marketplace, or self-hosted open-weight models.
  10. Vendor factors: stability, deprecation policies, support and rate limits.

Proprietary vs open-weight

Proprietary hosted models often lead on frontier capability and come with managed infrastructure, safety systems and features. Open-weight models (weights downloadable under a licence) offer control, the ability to self-host for data residency or privacy, and fine-tuning flexibility, but require infrastructure, security and operational expertise. Check licences carefully; "open" licences vary in commercial terms.

Many organisations use both: hosted frontier models for hard tasks, smaller or self-hosted models for high-volume or sensitive workloads.

Matching tasks to model types

TaskWhat to prioritise
Customer chatLatency, tone, tool calling, safety, languages
Document extractionVision and PDF handling, structured outputs, cost per page
Research and analysisReasoning, long context, citations
Coding assistanceReasoning, tool use, code quality on your stack
Bulk classificationCost, throughput, batch support; small models often suffice
Voice agentsStreaming STT/TTS quality, end-to-end latency
Image/video creationPrompt adherence, style control, licensing terms

The portfolio approach

Rather than one model for everything, many teams run a portfolio:

  • A default capable model for most tasks.
  • A fast, cheap model for high-volume simple steps.
  • A reasoning configuration for hard cases.
  • Specialised models where needed (embeddings, STT, image generation).

An abstraction layer in your code (a single internal interface that can call different providers) makes switching easier as the market changes. Keep prompts and evaluation sets per task so switching is a measured decision.

Worked example

A marketing agency serving clients in Pakistan, the UAE and the UK chooses models for four workflows:

  • Social caption drafting in English, Urdu and Arabic: they test three models on 30 real briefs per language; one clearly leads on Urdu, another on Arabic, so they route by language.
  • Invoice extraction: a vision-capable model with structured outputs; cost per page is the deciding factor between two similar performers.
  • Campaign analysis: a reasoning model for quarterly deep dives only.
  • Comment moderation: a small, fast model with escalation of uncertain cases.

Traps

  • Choosing by leaderboard rank or brand alone.
  • Ignoring total cost (retries, reasoning tokens, human review).
  • Forgetting data terms until procurement blocks deployment.
  • Locking into one provider's proprietary features without an exit plan.

How to shortlist in a fast-moving market

The major closed-model families in 2026 are Anthropic's Claude (Opus, Sonnet and Haiku tiers), OpenAI's GPT models, and Google's Gemini (Pro, Flash and Flash-Lite tiers), alongside strong open-weight families from several labs. Names and versions change every few months, so build your shortlist from the providers' current model pages and the Models APIs, which report context windows and capabilities programmatically.

A practical shortlisting sequence:

  1. Hard filters first: required modalities (PDF, audio, video in; images or speech out), context length, data residency and contractual terms, platform availability (direct API, AWS, Google Cloud, Azure).
  2. Tier by task: a flagship model for hard reasoning, a mid-tier model for the bulk of work, a fast tier for high-volume simple steps.
  3. Two or three candidates per task into your private benchmark (next lesson).

Hands-on: list models and capabilities programmatically

import anthropic
from openai import OpenAI

for m in anthropic.Anthropic().models.list():
    print("anthropic", m.id, getattr(m, "max_input_tokens", None))

for m in OpenAI().models.list():
    print("openai", m.id)

Use this in a scheduled job that alerts you when new models appear or old ones disappear, and cross-check deprecation pages. The Gemini API offers a similar models listing in its SDK.

Multimodal capability matrix (fill in for your shortlist)

Capability            | Candidate A | Candidate B | Candidate C
----------------------|-------------|-------------|------------
Image input           |             |             |
PDF input + citations |             |             |
Audio input           |             |             |
Video input           |             |             |
Image output          |             |             |
Structured outputs    |             |             |
Effort / thinking ctl |             |             |
Computer use          |             |             |
Data residency        |             |             |

Fill every cell from official documentation, with the date you checked. That table, plus your benchmark scores, is a decision record you can defend.

Going further

Maintain a simple model register: which models are used for which tasks, versions pinned, evaluation scores, costs, data terms and deprecation dates. Review it quarterly; this is how you keep pace with a fast-moving market without constant disruption.

Key takeaways

  • There is no single best model; choose by capability fit, modalities, context, latency, cost, reliability features, languages, compliance and deployment.
  • Proprietary hosted models and open-weight models each have trade-offs; many teams use both.
  • Run a portfolio: default, fast/cheap, reasoning and specialised models, behind an abstraction layer.
  • Avoid leaderboard-only decisions; keep a model register with scores, costs, data terms and deprecations.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which factor is most often overlooked until procurement blocks deployment?
  2. For bulk classification of 1 million short texts, what should you prioritise?
  3. Why use an abstraction layer for model calls in your code?

Put it into practice

List three AI tasks in your organisation. For each, write the top three selection criteria from this lesson and one candidate model type to test.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.