---
title: "Choosing models and an architecture that can change"
description: "The model is a component, not the product Model choice matters, but it changes: providers release new models several times a year, retire old ones, and…"
url: https://optimizeall.com/learn/building-ai-products-and-workflows/choosing-models-and-architecture
updated: 2026-10-05
---

Building AI Products & Workflows · Architecture, models and orchestration · lesson 6 of 18 · 14 min

# Choosing models and an architecture that can change

## The model is a component, not the product

Model choice matters, but it changes: providers release new models several times a year, retire old ones, and change prices. Teams that hard-wire one model into every prompt and code path pay for it at every change. Teams that treat models as **replaceable components behind a thin layer**, chosen per task and verified by evaluation, can adopt improvements in days.

## The main options (at the time of writing)

| Option | Examples | Strengths | Trade-offs |
|---|---|---|---|
| Frontier proprietary models via first-party APIs | Anthropic Claude, OpenAI GPT, Google Gemini families | Highest capability on hard reasoning, coding and agentic tasks; rich features (tools, structured outputs, caching, batch) | Per-token cost; data leaves your environment under the provider's terms |
| The same models via cloud platforms | Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry | Existing cloud contracts, regional hosting, private networking and IAM | Feature availability can lag or differ; check per feature |
| Smaller, faster models | "Mini", "flash" or "haiku"-class tiers | Low latency and cost for classification, extraction, routing | Weaker on complex reasoning |
| Open-weight models | Families such as Llama, Mistral, Qwen, Gemma, DeepSeek and OpenAI's gpt-oss | Self-hosting, data control, fine-tuning, predictable cost at scale | You run infrastructure, safety and upgrades; licences vary |

Names and line-ups change quickly; check each provider's current model list and deprecation schedule.

## How to choose, per task

1. **Define the task and its quality bar** (from your AI brief and evaluation set).
2. **Shortlist 2 to 4 models** across tiers, including one cheaper option.
3. **Run your evaluation set** on each, recording quality by tag, latency (median and 95th percentile), cost per successful outcome and failure modes.
4. **Tune the settings that matter.** Many current models expose a reasoning or "effort" control (for example Anthropic's `effort` setting or other providers' reasoning-effort parameters). Lower effort often holds quality on simple routes at much lower cost and latency; higher effort pays off on hard ones. Measure rather than guess.
5. **Consider data and regional constraints**: where processing happens, retention, and whether a cloud platform or self-hosted option is required.
6. **Decide per route**, not globally. A support product might use a small model to classify, a frontier model at medium effort to draft, and a batch job overnight for summaries.

## Architecture that survives model changes

```text
Product features
   |
Thin AI service layer  (your code: prompts, tools, schemas, routing, retries, logging)
   |
Model gateway          (keys, budgets, rate limits, provider fallbacks, usage tracking)
   |
Providers / cloud platforms / self-hosted models
```

- **Keep prompts, tool definitions and schemas in version-controlled files**, not scattered in code.
- **Route by configuration**, so switching a route's model is a config change plus an evaluation run.
- **A gateway** (open-source gateways such as LiteLLM, managed routers, or your cloud's AI gateway) centralises keys, budgets, logging and fallbacks. Keep it thin; avoid gateways that hide provider-specific features you need.
- **Pin and record versions** in every log line, and read providers' deprecation notices.
- **Re-run evaluations before any model switch**, even an "upgrade". Newer models can change tone, length, formatting or tool-calling habits.

## Hands-on: config-driven routing with fallback and effort

```yaml
# routes.yaml
classify_ticket: {provider: anthropic, model: claude-haiku-4-5, max_tokens: 200}
draft_reply:     {provider: anthropic, model: claude-opus-5, effort: medium, max_tokens: 2000,
                  fallback: {provider: anthropic, model: claude-sonnet-5, max_tokens: 2000}}
```

```python
import os, yaml
import anthropic

ROUTES = yaml.safe_load(open("routes.yaml"))
client = anthropic.Anthropic()          # one client per provider in a real gateway

def call(route: str, system: str, user: str) -> dict:
    cfg = ROUTES[route]
    for attempt_cfg in (cfg, cfg.get("fallback")):
        if not attempt_cfg:
            break
        kwargs = dict(model=attempt_cfg["model"], max_tokens=attempt_cfg["max_tokens"],
                      system=system, messages=[{"role": "user", "content": user}])
        if "effort" in attempt_cfg:
            kwargs["output_config"] = {"effort": attempt_cfg["effort"]}   # supported on current models; check docs
        try:
            resp = client.messages.create(**kwargs)
            return {"text": "".join(b.text for b in resp.content if b.type == "text"),
                    "model": attempt_cfg["model"], "usage": resp.usage, "stop": resp.stop_reason}
        except (anthropic.RateLimitError, anthropic.InternalServerError, anthropic.APIConnectionError):
            continue                                   # try the fallback route
    raise RuntimeError(f"route {route} failed on primary and fallback")
```

Model IDs above are examples current at the time of writing; always take them from the provider's model list. The flagship **Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API** course covers multi-provider integration in depth.

## Worked example

A UAE travel company's support copilot initially used one frontier model at maximum settings for everything. Evaluation by route showed classification quality was identical on a small model, and drafting quality held at medium effort for routine requests. After routing (small model to classify, frontier model at medium effort to draft, high effort only for complaint escalations), median latency and cost per resolved ticket both fell substantially, with no drop in quality scores. When the provider later released a new model, the team switched one route, re-ran the suite, and rolled it out in a day.

## Go deeper

Go deeper: **Open-Weight and Local AI: Run, Choose and Deploy Your Own Models** and **Fine-Tuning, Distillation and Custom Models** cover self-hosted and customised models; **Multimodal & Reasoning Models in Practice** covers reasoning and multimodal capabilities.

## Video lecture: Choosing models and an architecture that can change

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Choosing models
2. Analogy: staffing a team
3. The options
4. Choosing per task
5. Constraints and routes
6. Architecture
7. Simple example: two routes
8. Worked example: travel support copilot
9. Business example (illustrative)
10. Hands-on in the lesson
11. Common mistakes
12. How you'll know it's right
13. Watch me do it: config routing
14. Recap
15. Try this now (30 minutes)

## Lecture transcript

### Choosing models

The model you pick today will not be the model you use next year. Providers release new models several times a year, retire old ones and change prices. If your product has one model name hard-wired into every prompt and code path, every change becomes a project. In this lesson you'll learn the main model options and their trade-offs, how to choose per task using your own evaluation, and an architecture that lets you switch models in a day instead of a quarter.

### Analogy: staffing a team

Here's an analogy. Choosing models is like staffing a team. You wouldn't hire a senior partner to sort the post, or an intern to negotiate a merger. You match the person to the task, and you organise the office so people can change roles without rebuilding the building. Small models sort the post, frontier models handle the hard thinking, and your architecture lets you reassign them easily.

### The options

Here are the main options. Frontier models from first-party APIs, like the Claude, GPT and Gemini families, offer the highest capability on hard reasoning, coding and agentic work, plus rich features such as tools, structured outputs, caching and batch processing. The same models are often available through cloud platforms, like Amazon Bedrock, Google Cloud's Vertex AI and Microsoft Foundry, which helps with existing contracts, regional hosting and private networking, though feature availability can differ. Smaller, faster tiers are cheap and quick for classification, extraction and routing. And open-weight models, from families like Llama, Mistral, Qwen, Gemma, DeepSeek and gpt-oss, give you self-hosting and data control, at the cost of running infrastructure and safety yourself.

### Choosing per task

Choose per task, not per company. Define the task's quality bar from your brief and evaluation set. Shortlist two to four models across tiers, including a cheaper one. Run your evaluation set on each and record quality by tag, median and slow-tail latency, cost per successful outcome and failure modes. Then tune the settings that matter. Many current models expose an effort or reasoning control. Lower effort often holds quality on simple routes at much lower cost and latency. Higher effort pays off on hard ones. Measure, don't guess.

### Constraints and routes

Don't forget data and regional constraints. Where is data processed? For how long is it retained? Do you need a cloud platform in a specific region, or self-hosting? A bank in Riyadh and a creator tool in London may make different choices for the same task. And decide per route. A support product might use a small model to classify, a frontier model at medium effort to draft, and an overnight batch job for summaries.

### Architecture

Now the architecture. Your product features call a thin AI service layer that you own: prompts, tools, schemas, routing, retries and logging. Beneath it sits a model gateway handling keys, budgets, rate limits, fallbacks and usage tracking. Beneath that, providers, cloud platforms or self-hosted models. Keep prompts and schemas in version-controlled files. Route by configuration so switching a model is a config change plus an evaluation run. Pin and record versions in every log. And re-run evaluations before any switch, even an upgrade, because newer models can change tone, length, formatting and tool habits.

### Simple example: two routes

A simple example. A creator agency has two AI routes. One tags incoming brand emails into five categories. The other drafts detailed campaign proposals. They test a small, fast model and a frontier model on fifty emails and ten briefs. For tagging, both score the same, and the small one is far cheaper and faster. For proposals, the frontier model is clearly better. Two routes, two models, decided in an afternoon.

### Worked example: travel support copilot

Here's a worked example. A UAE travel company's support copilot used one frontier model at maximum settings for everything. Evaluation by route showed classification was just as good on a small model, and drafting held quality at medium effort for routine requests. After routing, with high effort only for complaint escalations, median latency and cost per resolved ticket both fell substantially with no drop in quality scores. When the provider later released a new model, they switched one route, re-ran the suite and rolled it out in a day.

### Business example (illustrative)

Illustrative numbers for the travel copilot. Before routing, median latency was about nine seconds and cost per resolved ticket was the biggest line in the AI budget. After routing, classification took under a second, routine drafts about four seconds, and cost per resolved ticket fell by more than half. Quality scores on four hundred test tickets stayed within a point. When the new model arrived, one config line changed and the suite ran overnight.

### Hands-on in the lesson

The hands-on section shows a routes file that maps each task to a provider, model, effort level and fallback, and a short Python function that calls the primary route, falls back on rate limits, server errors or connection failures, and returns which model actually answered, along with usage. Model IDs are examples current at the time of writing, so always take them from the provider's model list.

### Common mistakes

Common mistakes. Using one model everywhere at maximum settings. Choosing by public benchmark instead of your own evaluation. Hard-coding model names across the codebase. Ignoring where data is processed. Treating newer as automatically better. And forgetting to log which model and version produced each output, so you can't explain a change in behaviour when it happens.

### How you'll know it's right

How will you know your model choices and architecture are right? Each route has a named model, a documented reason and evaluation results. Switching a route takes a configuration change and an evaluation run, not a code change. Cost per successful outcome is tracked per route. And when a provider announces a deprecation, you know within minutes which routes are affected.

### Watch me do it: config routing

Watch me do it. I open routes dot yaml. Classify ticket uses a small model with two hundred max tokens. Draft reply uses a frontier model at medium effort, with a fallback to another model. Next, the call function loads the route config. For the primary and then the fallback, it builds the request, adds output config effort if the route sets one, and calls the API. On rate-limit, server or connection errors it moves to the fallback; any other error surfaces immediately. It returns the text, the model that actually answered, usage and the stop reason. I run the draft route while simulating a rate-limit on the primary, and the result shows the fallback model's name, which I log. Finally, I change classify ticket to a different small model in the YAML, run the evaluation suite on that route only, and compare accuracy and latency before merging.

### Recap

To recap: models are components. Know the options, choose per route using your own evaluation and effort settings, respect data constraints, and build a thin service layer and gateway with versioned prompts and config routing. Your next step is to list the AI routes in one product or workflow, shortlist two models for each, and write a routes entry with a fallback. For deeper integration work, see the Integrating AI Platforms course. Next, orchestrating multi-step workflows.

### Try this now (30 minutes)

Try this now. List every place your product or workflow calls an AI model. Group them into routes, like classify, extract, draft and summarise. For each route, write the current model, the quality bar and one cheaper candidate. Then write a routes file entry for the two most expensive routes, with a fallback, and schedule a short evaluation to test the cheaper candidate.

## Key takeaways

- Treat models as replaceable components chosen per route and verified by your own evaluation.
- Options span frontier APIs, the same models via cloud platforms, smaller fast tiers and open-weight models, each with trade-offs.
- Tune effort or reasoning settings and measure quality, latency and cost per successful outcome before choosing.
- A thin AI service layer plus a gateway, versioned prompts and config-driven routing let you switch models in days.

## Try it

List the AI routes in one product or workflow. For each, shortlist two models, define the quality bar and write a routes.yaml entry with a fallback.

- [Previous: Build vs buy (and blend)](https://optimizeall.com/learn/building-ai-products-and-workflows/build-vs-buy)
- [Next: Workflow orchestration patterns for AI features](https://optimizeall.com/learn/building-ai-products-and-workflows/workflow-orchestration-patterns)
- [All lessons of Building AI Products & Workflows](https://optimizeall.com/learn/building-ai-products-and-workflows)
