---
title: "Model routing: the right model for each request"
description: "What routing is Routing chooses which model (and settings) handles each request. Fallbacks (lesson 13) react to failures; routing proactively matches…"
url: https://optimizeall.com/learn/ai-platform-apis-integration/model-routing-strategies
updated: 2026-10-05
---

Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Multi-provider architecture · lesson 14 of 19 · 14 min

# Model routing: the right model for each request

## What routing is

**Routing** chooses which model (and settings) handles each request. Fallbacks (lesson 13) react to failures; routing proactively matches work to capability and cost. Done well, it cuts cost substantially while holding quality; done badly, it adds complexity, fragments caches and degrades hard cases.

## Routing strategies

| Strategy | How it decides | Good for |
|---|---|---|
| **Static by task** | Each feature has a configured model | Most products; simple and predictable |
| **Rule-based** | Input length, language, customer tier, presence of attachments | Clear, explainable differences |
| **Classifier-based** | A small model or classifier labels difficulty/intent first | High-volume traffic with a mix of easy and hard |
| **Cascade** | Try a cheap model; escalate if a check fails (validation, confidence, judge) | Tasks with an automatic quality check |
| **Effort routing** | Same model, different reasoning effort per request | Keeps one cache namespace; often simplest win |

Start with **static by task**. Move to rules or classifiers only when your cost data shows a large volume of easy requests being served by an expensive model.

## Measure the simple alternative first

Before building a cascade across models, test the strongest model at **lower effort** or a mid-tier model for the whole task. With current reasoning models, lower effort often matches older top-tier quality at a fraction of the cost, and keeping one model preserves prompt-cache hits. Judge by **cost per completed task**, including retries and escalations.

## Designing a cascade

1. **Cheap attempt** with a small model.
2. **Automatic check**: schema validation, business rules, a confidence field, or a lightweight judge.
3. **Escalate** to a stronger model when the check fails, passing the original input (not the failed output, unless useful).
4. **Log** which path each request took; review escalations.

Cascades only work when step 2 is reliable. If you can't detect bad outputs automatically, route by input characteristics instead.

## Hands-on: a rule-plus-classifier router

```python
import re

ROUTES = {   # loaded from config
    "simple":  ("google", "gemini-flash-latest"),
    "default": ("anthropic", "claude-sonnet-5"),
    "complex": ("anthropic", "claude-opus-5"),
}
ARABIC_OR_URDU = re.compile(r"[؀-ۿ]")

def classify_difficulty(text: str) -> str:
    """Cheap heuristic first; replace with a small-model classifier if needed."""
    if len(text) < 200 and "?" in text and not re.search(r"\b(contract|legal|refund|complaint)\b", text, re.I):
        return "simple"
    if len(text) > 4000 or re.search(r"\b(analy[sz]e|compare|strategy|forecast)\b", text, re.I):
        return "complex"
    return "default"

def route(text: str, tenant_tier: str) -> tuple[str, str]:
    level = classify_difficulty(text)
    if ARABIC_OR_URDU.search(text) and level == "simple":
        level = "default"            # our eval showed the small model is weaker in Arabic/Urdu
    if tenant_tier == "enterprise" and level == "simple":
        level = "default"            # contractual quality commitment
    return ROUTES[level]
```

Evaluate the router itself: label 200 real requests with the cheapest model that produces an acceptable answer, then measure how often the router picks a model that is too weak (quality risk) or too strong (cost waste).

## Worked example: a support assistant for a PK telecom reseller

Traffic: 60% simple balance/package questions, 30% troubleshooting, 10% complaints and billing disputes (illustrative). Routing: simple questions in English go to a small fast model; Urdu and Roman-Urdu messages and troubleshooting go to a mid-tier model; complaints and disputes go to the strongest model with a human-handover rule. The team first tested "mid-tier for everything at low effort" as a baseline; routing beat it on cost with equal quality, so they kept it, but reviewed escalations weekly.

## Operating a router over time

- **Re-evaluate on every model launch**: a new small model may absorb traffic that needed a mid-tier model last quarter.
- **Watch drift in the traffic mix**: a marketing campaign or new market can shift the share of hard requests; alert when route proportions move sharply.
- **Keep conversations sticky**: once a conversation starts on a model, keep it there unless there is a strong reason to escalate, for consistent tone and cache reuse.
- **Make routes explainable**: log the rule or classifier score behind each decision so support teams can answer "why did this customer get a worse answer?"

## Pitfalls

- Building a complex router before measuring the one-model baseline.
- Cascades without a reliable automatic quality check.
- Routers that ignore language, which silently lowers quality for some customers.
- Losing cache efficiency by spreading one conversation across models.

## Measuring success

Cost per completed task versus the baseline, quality per route (and per language), under-routing rate (too-weak model chosen), escalation rate, and cache hit ratio.

## Video lecture: Model routing: the right model for each request

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Model routing
2. Why it matters
3. Routing strategies
4. Measure the simple baseline first
5. Cascades
6. Simple example: the bookshop
7. Hands-on router
8. Business example: PK telecom support
9. Operating a router
10. Evaluate the router
11. Common mistakes
12. Effort routing first
13. Deeper: telecom routing (illustrative)
14. Watch me do it: route()
15. Recap + try this now

## Lecture transcript

### Model routing

Sending every request to your most powerful model is like sending a senior partner to answer every phone call. The quality is great, but most calls didn't need them, and the bill is huge. Sending everything to the cheapest model is the opposite mistake. Model routing is the art of sending each request to the right model. In this lesson you'll learn the main routing strategies, when they're worth it, and how to test them honestly.

### Why it matters

Why does this matter? In most products, the mix of requests is lopsided: lots of easy ones, some medium ones, and a few hard ones. If you pay top tier prices for the easy majority, costs balloon. If you send the hard minority to a weak model, quality suffers exactly where it matters most. Think of a hospital triage desk: minor cuts, broken arms and emergencies each go to the right team. That's routing.

### Routing strategies

Here are the strategies. Static by task: each feature has a configured model, simple and predictable, and right for most products. Rule based: route by input length, language, customer tier or attachments. Classifier based: a small model labels difficulty or intent first. Cascade: try a cheap model, check the output automatically, and escalate if it fails. And effort routing: the same model with different reasoning effort per request, which keeps one cache and is often the simplest win.

### Measure the simple baseline first

Here's the key idea that saves teams months: measure the simple alternative first. Before building a cascade across several models, test the strongest model at a lower effort setting, or a mid tier model for the whole task. Current reasoning models at lower effort often match older top tier quality for much less, and one model keeps your prompt cache working. Compare everything by cost per completed task, including retries and escalations.

### Cascades

How do cascades work? First, a cheap attempt with a small model. Second, an automatic check: schema validation, business rules, a confidence field or a lightweight judge. Third, escalate to a stronger model when the check fails, sending the original input. Fourth, log which path each request took, and review escalations. The whole thing depends on step two. If you can't reliably detect a bad output automatically, route by input characteristics instead.

### Simple example: the bookshop

A simple example. An online bookshop's assistant gets two kinds of questions: do you have this title in stock, and can you recommend a book for my grandmother who loved a particular novel. Stock questions are simple lookups, so a small fast model plus a tool handles them. Recommendations need taste and nuance, so they go to a stronger model. A single keyword rule based on intent covers almost all traffic. Nothing fancy, but it roughly halves costs without anyone noticing a difference.

### Hands-on router

The lesson's code shows a router combining rules and a cheap heuristic classifier. Short questions without sensitive words are simple, long or analytical requests are complex, and everything else is default. Two business rules then adjust it: Arabic or Urdu text never goes to the small model, because the team's eval showed it was weaker there, and enterprise customers never get the small model, because of a contractual quality commitment. Routes themselves come from configuration.

### Business example: PK telecom support

Now a realistic business example, with illustrative numbers. A telecom reseller in Pakistan sees about sixty percent simple balance and package questions, thirty percent troubleshooting, and ten percent complaints and billing disputes. Simple English questions go to a small fast model. Urdu and Roman Urdu messages, and troubleshooting, go to a mid tier model. Complaints and disputes go to the strongest model with a human handover rule. Their baseline, mid tier for everything at low effort, was good, but routing beat it on cost at equal quality, so they kept it and review escalations weekly.

### Operating a router

Routers need looking after, like any production system. Re evaluate every time a new model launches, because a new small model might absorb traffic that needed a mid tier model last quarter. Watch the traffic mix too: a marketing campaign or a new market can suddenly raise the share of hard requests. Keep a conversation on one model once it starts, for consistent tone and cache reuse. And log the reason for every routing decision, so support can answer, why did this customer get a weaker answer?

### Evaluate the router

How do you evaluate the router itself? Label a few hundred real requests with the cheapest model that gives an acceptable answer. Then measure how often your router picks a model that's too weak, which risks quality, or too strong, which wastes money. Break results down by language and customer tier, because averages hide the groups you're underserving.

### Common mistakes

Common mistakes. Building a complex router before measuring the one model baseline. Cascades without a reliable automatic check. Ignoring language, which silently lowers quality for some customers. And spreading one conversation across models, which loses cache efficiency and can make the assistant's tone inconsistent.

### Effort routing first

One more idea that's easy to overlook: routing by effort within a single model. Instead of sending easy questions to a different, smaller model, you send them to the same strong model with a low reasoning effort, and hard ones with high effort. You keep one prompt style, one cache and one set of behaviors to test, while still saving money on the easy majority. For many teams, this is the best first step before any multi model routing at all.

### Deeper: telecom routing (illustrative)

Let's deepen the Pakistani telecom reseller example with illustrative numbers. About sixty thousand messages a month. The baseline, one mid tier model at low effort for everything, was good and simple. The router sent English balance questions to a small model, Urdu, Roman Urdu and troubleshooting to the mid tier model, and disputes to the strongest model with a human handover. Cost per resolved conversation fell by a meaningful margin compared with the baseline, and quality held across all three languages in the eval. The weekly escalation review found one recurring issue: Roman Urdu balance questions misclassified as simple, so they tightened the language rule.

### Watch me do it: route()

Watch me do it. Let's walk through the router function. Routes maps three levels to a provider and model: simple, default and complex. A regular expression detects Arabic script characters, which also covers Urdu. Classify difficulty returns simple for short questions without sensitive words like contract, refund or complaint; complex for very long inputs or analytical words like compare or forecast; otherwise default. Route calls the classifier, then applies two business rules. If the text contains Arabic script and the level is simple, it bumps it to default, because the eval showed the small model is weaker there. If the tenant is enterprise and the level is simple, it also bumps to default, honoring a contract. Finally it returns the provider and model for that level. I test four messages: an English balance question goes small, an Urdu one goes default, a refund complaint goes default, and a forecast request goes complex.

### Recap + try this now

Quick recap. Routing matches requests to models proactively. Start static by task, measure a simple baseline, add rules, classifiers or cascades only when data justifies them, make cascade checks reliable, and evaluate the router by language and tier. Try this now: label a hundred real requests with the cheapest acceptable model, compare a one model baseline against a simple router on cost and quality, and decide with evidence whether routing is worth it.

## Key takeaways

- Routing proactively matches each request to a model and settings; fallbacks react to failures.
- Start with static routing by task; add rules, classifiers or cascades only when cost data justifies it.
- Always test the simple baseline first, such as one model at lower effort.
- Cascades need a reliable automatic quality check before escalation.
- Evaluate the router itself, including per-language quality.

## Try it

Label 100 real requests with the cheapest acceptable model, compare a one-model baseline with a simple router on cost and quality, and decide whether routing is worth it.

- [Previous: Multi-provider abstraction and fallbacks](https://optimizeall.com/learn/ai-platform-apis-integration/provider-abstraction-and-fallbacks)
- [Next: Cloud AI platforms: Amazon Bedrock, Microsoft Foundry and Gemini Enterprise Agent Platform](https://optimizeall.com/learn/ai-platform-apis-integration/cloud-ai-platforms)
- [All lessons of Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API](https://optimizeall.com/learn/ai-platform-apis-integration)
