Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APIMulti-provider architecture · Lesson 14 of 19

Model routing: the right model for each request

Article · 14 min · 9 min lecture

Video lecture

Model routing: the right model for each request

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Model routing

  • Right model per request
  • Routing strategies
  • Baselines and cascades
  • Measuring the router

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What routing is

Routing chooses which model (and settings) handles each request. Fallbacks (lesson 13) react to failures; routing proactively matches work to capability and cost. Done well, it cuts cost substantially while holding quality; done badly, it adds complexity, fragments caches and degrades hard cases.

Routing strategies

StrategyHow it decidesGood for
Static by taskEach feature has a configured modelMost products; simple and predictable
Rule-basedInput length, language, customer tier, presence of attachmentsClear, explainable differences
Classifier-basedA small model or classifier labels difficulty/intent firstHigh-volume traffic with a mix of easy and hard
CascadeTry a cheap model; escalate if a check fails (validation, confidence, judge)Tasks with an automatic quality check
Effort routingSame model, different reasoning effort per requestKeeps one cache namespace; often simplest win

Start with static by task. Move to rules or classifiers only when your cost data shows a large volume of easy requests being served by an expensive model.

Measure the simple alternative first

Before building a cascade across models, test the strongest model at lower effort or a mid-tier model for the whole task. With current reasoning models, lower effort often matches older top-tier quality at a fraction of the cost, and keeping one model preserves prompt-cache hits. Judge by cost per completed task, including retries and escalations.

Designing a cascade

  1. Cheap attempt with a small model.
  2. Automatic check: schema validation, business rules, a confidence field, or a lightweight judge.
  3. Escalate to a stronger model when the check fails, passing the original input (not the failed output, unless useful).
  4. Log which path each request took; review escalations.

Cascades only work when step 2 is reliable. If you can't detect bad outputs automatically, route by input characteristics instead.

Hands-on: a rule-plus-classifier router

import re

ROUTES = {   # loaded from config
    "simple":  ("google", "gemini-flash-latest"),
    "default": ("anthropic", "claude-sonnet-5"),
    "complex": ("anthropic", "claude-opus-5"),
}
ARABIC_OR_URDU = re.compile(r"[؀-ۿ]")

def classify_difficulty(text: str) -> str:
    """Cheap heuristic first; replace with a small-model classifier if needed."""
    if len(text) < 200 and "?" in text and not re.search(r"\b(contract|legal|refund|complaint)\b", text, re.I):
        return "simple"
    if len(text) > 4000 or re.search(r"\b(analy[sz]e|compare|strategy|forecast)\b", text, re.I):
        return "complex"
    return "default"

def route(text: str, tenant_tier: str) -> tuple[str, str]:
    level = classify_difficulty(text)
    if ARABIC_OR_URDU.search(text) and level == "simple":
        level = "default"            # our eval showed the small model is weaker in Arabic/Urdu
    if tenant_tier == "enterprise" and level == "simple":
        level = "default"            # contractual quality commitment
    return ROUTES[level]

Evaluate the router itself: label 200 real requests with the cheapest model that produces an acceptable answer, then measure how often the router picks a model that is too weak (quality risk) or too strong (cost waste).

Worked example: a support assistant for a PK telecom reseller

Traffic: 60% simple balance/package questions, 30% troubleshooting, 10% complaints and billing disputes (illustrative). Routing: simple questions in English go to a small fast model; Urdu and Roman-Urdu messages and troubleshooting go to a mid-tier model; complaints and disputes go to the strongest model with a human-handover rule. The team first tested "mid-tier for everything at low effort" as a baseline; routing beat it on cost with equal quality, so they kept it, but reviewed escalations weekly.

Operating a router over time

  • Re-evaluate on every model launch: a new small model may absorb traffic that needed a mid-tier model last quarter.
  • Watch drift in the traffic mix: a marketing campaign or new market can shift the share of hard requests; alert when route proportions move sharply.
  • Keep conversations sticky: once a conversation starts on a model, keep it there unless there is a strong reason to escalate, for consistent tone and cache reuse.
  • Make routes explainable: log the rule or classifier score behind each decision so support teams can answer "why did this customer get a worse answer?"

Pitfalls

  • Building a complex router before measuring the one-model baseline.
  • Cascades without a reliable automatic quality check.
  • Routers that ignore language, which silently lowers quality for some customers.
  • Losing cache efficiency by spreading one conversation across models.

Measuring success

Cost per completed task versus the baseline, quality per route (and per language), under-routing rate (too-weak model chosen), escalation rate, and cache hit ratio.

Key takeaways

  • Routing proactively matches each request to a model and settings; fallbacks react to failures.
  • Start with static routing by task; add rules, classifiers or cascades only when cost data justifies it.
  • Always test the simple baseline first, such as one model at lower effort.
  • Cascades need a reliable automatic quality check before escalation.
  • Evaluate the router itself, including per-language quality.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Before building a multi-model cascade, what should you measure first?
  2. A cascade sends everything to a small model first. What makes it safe?
  3. Your router sends short questions to a small model, but Urdu questions get worse answers. What should you do?

Put it into practice

Label 100 real requests with the cheapest acceptable model, compare a one-model baseline with a simple router on cost and quality, and decide whether routing is worth it.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.