Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsHybrid routing and capstone · Lesson 15 of 16

Hybrid routing between local and API models

Article · 16 min · 8 min lecture

Video lecture

Hybrid routing between local and API models

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Hybrid routing

  • Local and cloud, per request
  • Four strategies
  • Sensitivity, confidence, proof

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The end state for most teams

Few organizations go "all local" or "all API". The winning pattern is a router: a small piece of software that decides, per request, which model handles it, based on data sensitivity, task difficulty, cost, latency and availability. Done well, you get privacy where it matters, frontier quality where it matters, and a lower bill overall.

Four routing strategies

StrategyRuleExample
Sensitivity-basedAnything containing regulated or confidential data stays localMessages with national ID numbers, health data or client contracts go to the on-prem model
Task-basedRoute by task typeClassification and extraction local; complex drafting and multi-step reasoning to a stronger API model
Cascade (try cheap first)Local model answers; escalate if validation fails or confidence is lowLocal extractor; if schema validation fails twice, escalate (only if the data is allowed to leave)
Availability fallbackIf the primary is down or slow, use a backupAPI outage → local model with a "reduced quality" notice

Combine them in a fixed order: sensitivity rules first (they are hard constraints), then task, then cascade and fallback.

Detecting sensitivity

Use layered, testable rules:

  1. Source metadata: requests from the HR system or patient portal are sensitive by definition.
  2. Pattern detection: regular expressions for identifiers (for example Pakistani CNIC format, Emirates ID format, IBANs, email addresses, phone numbers).
  3. A local classifier: a small local model that flags confidential content categories.
  4. User choice: a visible "keep this private" toggle that forces local processing.

Crucially, the sensitivity check itself must run locally; sending text to a cloud API to ask whether it is sensitive defeats the purpose.

Measuring "confidence"

Language models do not give reliable confidence out of the box. Practical signals:

  • Validation: did output parse and pass business rules?
  • Self-consistency: do two samples agree on the label?
  • Retrieval signals: did retrieval find strong matches, or nothing relevant?
  • Explicit abstention: the prompt allows "I don't know", which triggers escalation.

Hands-on: a minimal router

# router.py: sensitivity-first hybrid router (pip install openai)
import os, re
from openai import OpenAI

LOCAL = OpenAI(base_url=os.getenv("LOCAL_BASE_URL", "http://localhost:11434/v1"), api_key=os.getenv("LOCAL_API_KEY", "local"))
CLOUD = OpenAI(base_url=os.getenv("CLOUD_BASE_URL"), api_key=os.getenv("CLOUD_API_KEY"))  # any OpenAI-compatible provider or gateway
LOCAL_MODEL = os.getenv("LOCAL_MODEL", "qwen3:8b")
CLOUD_MODEL = os.getenv("CLOUD_MODEL")          # set to a model your provider offers; never hard-code keys

SENSITIVE = [
    re.compile(r"\b\d{5}-\d{7}-\d\b"),           # Pakistani CNIC format
    re.compile(r"\b784-\d{4}-\d{7}-\d\b"),        # Emirates ID format
    re.compile(r"\b[A-Z]{2}\d{2}[A-Z0-9]{11,30}\b"),  # IBAN-like
    re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+"),      # email address
]
HARD_TASKS = {"strategy_memo", "multi_step_analysis"}

def is_sensitive(text: str, source: str, force_private: bool) -> bool:
    return force_private or source in {"hr", "patient_portal"} or any(p.search(text) for p in SENSITIVE)

def ask(client, model, prompt, timeout=60):
    r = client.chat.completions.create(model=model, timeout=timeout, temperature=0.2,
                                       messages=[{"role": "user", "content": prompt}])
    return r.choices[0].message.content or ""

def route(prompt: str, task: str, source: str = "web", force_private: bool = False) -> dict:
    if is_sensitive(prompt, source, force_private):
        return {"route": "local:sensitive", "answer": ask(LOCAL, LOCAL_MODEL, prompt)}
    if task in HARD_TASKS and CLOUD_MODEL:
        try:
            return {"route": "cloud:hard_task", "answer": ask(CLOUD, CLOUD_MODEL, prompt)}
        except Exception as e:            # outage, rate limit, timeout
            return {"route": f"local:fallback({type(e).__name__})", "answer": ask(LOCAL, LOCAL_MODEL, prompt)}
    answer = ask(LOCAL, LOCAL_MODEL, prompt)
    if "i don't know" in answer.lower() and CLOUD_MODEL:     # cascade on explicit abstention
        return {"route": "cloud:escalated", "answer": ask(CLOUD, CLOUD_MODEL, prompt)}
    return {"route": "local:default", "answer": answer}

if __name__ == "__main__":
    print(route("Summarize the attached note for CNIC 42101-1234567-1", task="summary"))

Log the route for every request (without the sensitive content) so you can measure distribution, cost and quality per route.

Gateways such as LiteLLM and several API-management products implement model aliases, fallbacks, budgets and logging for you; the logic above is what you configure in them.

Economics of routing

Model the monthly bill as:

monthly cost ≈ Σ over routes ( requests_route × avg_tokens_route × price_per_token_route ) + fixed local infrastructure

If 70% of requests are simple and move to a local model you already run, and the remaining 30% go to a stronger API, total cost and data exposure both fall, while quality on hard tasks rises. Use your own traffic and current price sheets; figures in vendor examples rarely match your mix.

Worked example: an agency in Lahore serving UK and Gulf clients

  • Client-confidential briefs (source metadata = client workspace flagged "NDA"): local only.
  • Social captions and hashtag suggestions: local model.
  • Quarterly strategy decks for clients who approved cloud processing: stronger API model.
  • API outage: local fallback with a banner.

After a month (illustrative): 78% of requests local, 22% API; API spend down versus the previous all-API setup; no confidential brief left the office network; editors rate strategy drafts higher because the strongest model is reserved for them.

Pitfalls

  • Running the sensitivity check in the cloud.
  • Silent escalation of sensitive data during an outage. Sensitivity rules must beat fallback rules.
  • Different system prompts per route causing inconsistent tone; share prompts and test both routes.
  • Not logging routes, so you cannot prove where data went.

How to measure success

A route log showing the share of traffic per route, zero sensitive requests on cloud routes (tested with seeded canary data), cost per route, and quality scores per route from your evaluation set.

Key takeaways

  • Route per request on sensitivity, task difficulty, cost, latency and availability
  • Apply sensitivity rules first; they must override fallbacks
  • Detect sensitivity locally with metadata, patterns, a local classifier and a user toggle
  • Use validation, self-consistency, retrieval strength and abstention as confidence signals
  • Log the route for every request and test with canary data

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. The cloud API is down. A request contains an Emirates ID number. What should the router do?
  2. Why must the sensitivity check run locally?
  3. Which is a practical confidence signal for escalation?

Put it into practice

Adapt the router to three request types from your organization, seed 20 canary requests with fake IDs, and verify none reach the cloud route.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.