Building AI Products & WorkflowsArchitecture, models and orchestration · Lesson 6 of 18

Choosing models and an architecture that can change

Article · 14 min · 9 min lecture

Video lecture

Choosing models and an architecture that can change

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Choosing models

  • Models as replaceable components
  • Options and trade-offs
  • Architecture that can change

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The model is a component, not the product

Model choice matters, but it changes: providers release new models several times a year, retire old ones, and change prices. Teams that hard-wire one model into every prompt and code path pay for it at every change. Teams that treat models as replaceable components behind a thin layer, chosen per task and verified by evaluation, can adopt improvements in days.

The main options (at the time of writing)

OptionExamplesStrengthsTrade-offs
Frontier proprietary models via first-party APIsAnthropic Claude, OpenAI GPT, Google Gemini familiesHighest capability on hard reasoning, coding and agentic tasks; rich features (tools, structured outputs, caching, batch)Per-token cost; data leaves your environment under the provider's terms
The same models via cloud platformsAmazon Bedrock, Google Cloud Vertex AI, Microsoft FoundryExisting cloud contracts, regional hosting, private networking and IAMFeature availability can lag or differ; check per feature
Smaller, faster models"Mini", "flash" or "haiku"-class tiersLow latency and cost for classification, extraction, routingWeaker on complex reasoning
Open-weight modelsFamilies such as Llama, Mistral, Qwen, Gemma, DeepSeek and OpenAI's gpt-ossSelf-hosting, data control, fine-tuning, predictable cost at scaleYou run infrastructure, safety and upgrades; licences vary

Names and line-ups change quickly; check each provider's current model list and deprecation schedule.

How to choose, per task

  1. Define the task and its quality bar (from your AI brief and evaluation set).
  2. Shortlist 2 to 4 models across tiers, including one cheaper option.
  3. Run your evaluation set on each, recording quality by tag, latency (median and 95th percentile), cost per successful outcome and failure modes.
  4. Tune the settings that matter. Many current models expose a reasoning or "effort" control (for example Anthropic's effort setting or other providers' reasoning-effort parameters). Lower effort often holds quality on simple routes at much lower cost and latency; higher effort pays off on hard ones. Measure rather than guess.
  5. Consider data and regional constraints: where processing happens, retention, and whether a cloud platform or self-hosted option is required.
  6. Decide per route, not globally. A support product might use a small model to classify, a frontier model at medium effort to draft, and a batch job overnight for summaries.

Architecture that survives model changes

Product features
   |
Thin AI service layer  (your code: prompts, tools, schemas, routing, retries, logging)
   |
Model gateway          (keys, budgets, rate limits, provider fallbacks, usage tracking)
   |
Providers / cloud platforms / self-hosted models
  • Keep prompts, tool definitions and schemas in version-controlled files, not scattered in code.
  • Route by configuration, so switching a route's model is a config change plus an evaluation run.
  • A gateway (open-source gateways such as LiteLLM, managed routers, or your cloud's AI gateway) centralises keys, budgets, logging and fallbacks. Keep it thin; avoid gateways that hide provider-specific features you need.
  • Pin and record versions in every log line, and read providers' deprecation notices.
  • Re-run evaluations before any model switch, even an "upgrade". Newer models can change tone, length, formatting or tool-calling habits.

Hands-on: config-driven routing with fallback and effort

# routes.yaml
classify_ticket: {provider: anthropic, model: claude-haiku-4-5, max_tokens: 200}
draft_reply:     {provider: anthropic, model: claude-opus-5, effort: medium, max_tokens: 2000,
                  fallback: {provider: anthropic, model: claude-sonnet-5, max_tokens: 2000}}
import os, yaml
import anthropic

ROUTES = yaml.safe_load(open("routes.yaml"))
client = anthropic.Anthropic()          # one client per provider in a real gateway

def call(route: str, system: str, user: str) -> dict:
    cfg = ROUTES[route]
    for attempt_cfg in (cfg, cfg.get("fallback")):
        if not attempt_cfg:
            break
        kwargs = dict(model=attempt_cfg["model"], max_tokens=attempt_cfg["max_tokens"],
                      system=system, messages=[{"role": "user", "content": user}])
        if "effort" in attempt_cfg:
            kwargs["output_config"] = {"effort": attempt_cfg["effort"]}   # supported on current models; check docs
        try:
            resp = client.messages.create(**kwargs)
            return {"text": "".join(b.text for b in resp.content if b.type == "text"),
                    "model": attempt_cfg["model"], "usage": resp.usage, "stop": resp.stop_reason}
        except (anthropic.RateLimitError, anthropic.InternalServerError, anthropic.APIConnectionError):
            continue                                   # try the fallback route
    raise RuntimeError(f"route {route} failed on primary and fallback")

Model IDs above are examples current at the time of writing; always take them from the provider's model list. The flagship Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API course covers multi-provider integration in depth.

Worked example

A UAE travel company's support copilot initially used one frontier model at maximum settings for everything. Evaluation by route showed classification quality was identical on a small model, and drafting quality held at medium effort for routine requests. After routing (small model to classify, frontier model at medium effort to draft, high effort only for complaint escalations), median latency and cost per resolved ticket both fell substantially, with no drop in quality scores. When the provider later released a new model, the team switched one route, re-ran the suite, and rolled it out in a day.

Go deeper

Go deeper: Open-Weight and Local AI: Run, Choose and Deploy Your Own Models and Fine-Tuning, Distillation and Custom Models cover self-hosted and customised models; Multimodal & Reasoning Models in Practice covers reasoning and multimodal capabilities.

Key takeaways

  • Treat models as replaceable components chosen per route and verified by your own evaluation.
  • Options span frontier APIs, the same models via cloud platforms, smaller fast tiers and open-weight models, each with trade-offs.
  • Tune effort or reasoning settings and measure quality, latency and cost per successful outcome before choosing.
  • A thin AI service layer plus a gateway, versioned prompts and config-driven routing let you switch models in days.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Evaluation shows a small model classifies tickets as well as a frontier model. What should you do?
  2. A provider releases a newer model. What is the safest way to adopt it?
  3. Why keep prompts, tool definitions and schemas in version-controlled files?

Put it into practice

List the AI routes in one product or workflow. For each, shortlist two models, define the quality bar and write a routes.yaml entry with a fallback.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.