Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsHybrid routing and capstone · Lesson 15 of 16
Hybrid routing between local and API models
Video lecture
Hybrid routing between local and API models
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Hybrid routing
Here is the question most teams eventually face: local or cloud? And the best answer is usually: both, per request. In this lesson you will build a router, a small piece of software that decides which model handles each request based on data sensitivity, task difficulty, cost, latency and availability. You will learn the four routing strategies, how to detect sensitive data locally, what counts as confidence, and how to prove where your data went.
0:33 Analogy: hospital triage
Here is an analogy for routing. Think of a hospital triage desk. Routine cases go to the general clinic, which is fast and inexpensive. Complex cases go to specialists. And some cases, like anything involving a child's records, follow special rules no matter how simple they look. A router is your triage desk for AI requests. The golden rule of triage applies too: the special rules are checked first, before anyone thinks about speed or cost.
1:06 Four strategies (in this order)
Four strategies. Sensitivity-based: anything with regulated or confidential data stays local. Task-based: simple jobs like classification and extraction go local; complex drafting and multi-step reasoning go to a stronger API model. Cascade: try the cheap local model first and escalate if validation fails or it abstains. And availability fallback: if the API is down, use the local model with a reduced-quality notice. Combine them in a fixed order. Sensitivity first, because it is a hard rule. Then task, then cascade and fallback.
1:42 Detecting sensitivity (locally)
How do you detect sensitive data? In layers. Source metadata: anything from the HR system or patient portal is sensitive by definition. Patterns: regular expressions for identifiers like Pakistani CNIC numbers, Emirates ID numbers, bank account numbers, emails and phone numbers. A small local classifier for confidential topics. And a visible keep-this-private toggle for users. The golden rule: the check itself must run locally. Sending text to the cloud to ask whether it is sensitive has already exposed it.
2:16 Confidence signals
What about confidence? Language models do not give you a trustworthy confidence number by default. Use observable signals instead. Did the output pass schema validation and business rules? Do two samples agree? Did retrieval find strong matches or nothing relevant? And did the model explicitly say it does not know, which your prompt should allow? Any of those can trigger escalation, if the data is allowed to leave.
2:46 Hands-on: router.py
The router in the lesson is about forty lines of Python. Two OpenAI-compatible clients, one local and one cloud, with keys and model names from environment variables. A sensitivity check with patterns and source rules. Hard tasks go to the cloud, with a local fallback if the cloud fails. Everything else goes local, escalating on explicit abstention. And every request logs its route, without the sensitive content, so you can measure and prove behavior. Gateways like LiteLLM can implement the same logic as configuration.
3:23 Routing economics
Let's talk money. Your monthly bill is the sum, across routes, of requests times average tokens times the price per token, plus your fixed local infrastructure. If most of your traffic is simple and moves to a local model you already run, and the hard minority goes to a stronger API, total cost and data exposure both fall while quality on the hard tasks improves. Use your own traffic mix and current price sheets. Vendor examples rarely match your reality.
3:58 Worked example: Lahore agency (illustrative)
A worked example. A Lahore agency serves UK and Gulf clients. Confidential briefs from workspaces flagged under an NDA stay local, always. Social captions go to the local model. Quarterly strategy decks, for clients who approved cloud processing, go to a stronger API model. During an API outage, requests fall back to local with a banner. After a month, in this illustrative case, about four in five requests run locally, API spend falls compared with the old all-API setup, and no confidential brief ever leaves the office network.
4:36 Simple example: two-person practice
A simple example. A two-person accounting practice in Manchester uses a local model for anything containing client names or bank details, and a cloud model for general questions like explain this new tax rule in plain English. Their router is just one rule: if the message contains a client name from their list or anything that looks like an account number, stay local. Everything else can go to the cloud. One rule, tested with ten fake examples, and they sleep better.
5:11 Routing UX
There is also a user experience side to routing. Tell people, simply, where their request was processed when it matters, for example a small label saying processed on our servers or processed by an external provider. Offer the keep-this-private toggle where users handle sensitive work. And when a fallback lowers quality, say so. Transparency builds trust, and it also supports the disclosure expectations that AI regulations and data protection rules increasingly set.
5:42 FAQ
A question that comes up in every routing project: will users notice when they are switched between models? Sometimes, in tone and style. To keep the experience consistent, share one system prompt and output format across routes, test both routes on the same evaluation set, and review a sample of each route's answers side by side. Another question: how often should routing rules change? Review them monthly with the route log in front of you. If a local model has improved, move more traffic to it. If a route's quality drops, tighten the rules.
6:23 Try this now
Try this now. Write your first three routing rules in plain English, in priority order. Rule one must be about sensitive data. Then write five fake requests that should trigger rule one, using made-up identifiers, and five that should not. Those ten requests are your first canary test set.
6:44 Watch me do it
Watch me do it. I open the router script and set the local base address, the cloud gateway address and the model names from environment variables. I add our organization's rule first: anything from the HR source stays local. Then I write twenty canary requests with fake identifiers: fake CNIC numbers, fake Emirates ID numbers, fake emails, mixed into ordinary questions. I run them all through the router and print only the route for each. Nineteen go local correctly. One slips to the cloud: an ID number written with spaces instead of dashes. I update the pattern to allow both, rerun, and all twenty stay local. Then I simulate a cloud outage by pointing the cloud address at nothing; hard tasks fall back to local with the right route label, and no sensitive canary leaves. Finally I check the route log shows counts per route without any request content.
7:49 Pitfalls + recap
Three pitfalls, then recap. Never run the sensitivity check in the cloud. Make sure sensitivity rules beat fallback rules, or an outage will quietly send private data out. And share prompts across routes so tone stays consistent. To recap: route per request, sensitivity first, observable confidence signals, and log every route. Your next step: adapt the router to three request types from your work, seed twenty fake-ID canary requests, and prove none reach the cloud.
The end state for most teams
Few organizations go "all local" or "all API". The winning pattern is a router: a small piece of software that decides, per request, which model handles it, based on data sensitivity, task difficulty, cost, latency and availability. Done well, you get privacy where it matters, frontier quality where it matters, and a lower bill overall.
Four routing strategies
| Strategy | Rule | Example |
|---|---|---|
| Sensitivity-based | Anything containing regulated or confidential data stays local | Messages with national ID numbers, health data or client contracts go to the on-prem model |
| Task-based | Route by task type | Classification and extraction local; complex drafting and multi-step reasoning to a stronger API model |
| Cascade (try cheap first) | Local model answers; escalate if validation fails or confidence is low | Local extractor; if schema validation fails twice, escalate (only if the data is allowed to leave) |
| Availability fallback | If the primary is down or slow, use a backup | API outage → local model with a "reduced quality" notice |
Combine them in a fixed order: sensitivity rules first (they are hard constraints), then task, then cascade and fallback.
Detecting sensitivity
Use layered, testable rules:
- Source metadata: requests from the HR system or patient portal are sensitive by definition.
- Pattern detection: regular expressions for identifiers (for example Pakistani CNIC format, Emirates ID format, IBANs, email addresses, phone numbers).
- A local classifier: a small local model that flags confidential content categories.
- User choice: a visible "keep this private" toggle that forces local processing.
Crucially, the sensitivity check itself must run locally; sending text to a cloud API to ask whether it is sensitive defeats the purpose.
Measuring "confidence"
Language models do not give reliable confidence out of the box. Practical signals:
- Validation: did output parse and pass business rules?
- Self-consistency: do two samples agree on the label?
- Retrieval signals: did retrieval find strong matches, or nothing relevant?
- Explicit abstention: the prompt allows "I don't know", which triggers escalation.
Hands-on: a minimal router
# router.py: sensitivity-first hybrid router (pip install openai)
import os, re
from openai import OpenAI
LOCAL = OpenAI(base_url=os.getenv("LOCAL_BASE_URL", "http://localhost:11434/v1"), api_key=os.getenv("LOCAL_API_KEY", "local"))
CLOUD = OpenAI(base_url=os.getenv("CLOUD_BASE_URL"), api_key=os.getenv("CLOUD_API_KEY")) # any OpenAI-compatible provider or gateway
LOCAL_MODEL = os.getenv("LOCAL_MODEL", "qwen3:8b")
CLOUD_MODEL = os.getenv("CLOUD_MODEL") # set to a model your provider offers; never hard-code keys
SENSITIVE = [
re.compile(r"\b\d{5}-\d{7}-\d\b"), # Pakistani CNIC format
re.compile(r"\b784-\d{4}-\d{7}-\d\b"), # Emirates ID format
re.compile(r"\b[A-Z]{2}\d{2}[A-Z0-9]{11,30}\b"), # IBAN-like
re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+"), # email address
]
HARD_TASKS = {"strategy_memo", "multi_step_analysis"}
def is_sensitive(text: str, source: str, force_private: bool) -> bool:
return force_private or source in {"hr", "patient_portal"} or any(p.search(text) for p in SENSITIVE)
def ask(client, model, prompt, timeout=60):
r = client.chat.completions.create(model=model, timeout=timeout, temperature=0.2,
messages=[{"role": "user", "content": prompt}])
return r.choices[0].message.content or ""
def route(prompt: str, task: str, source: str = "web", force_private: bool = False) -> dict:
if is_sensitive(prompt, source, force_private):
return {"route": "local:sensitive", "answer": ask(LOCAL, LOCAL_MODEL, prompt)}
if task in HARD_TASKS and CLOUD_MODEL:
try:
return {"route": "cloud:hard_task", "answer": ask(CLOUD, CLOUD_MODEL, prompt)}
except Exception as e: # outage, rate limit, timeout
return {"route": f"local:fallback({type(e).__name__})", "answer": ask(LOCAL, LOCAL_MODEL, prompt)}
answer = ask(LOCAL, LOCAL_MODEL, prompt)
if "i don't know" in answer.lower() and CLOUD_MODEL: # cascade on explicit abstention
return {"route": "cloud:escalated", "answer": ask(CLOUD, CLOUD_MODEL, prompt)}
return {"route": "local:default", "answer": answer}
if __name__ == "__main__":
print(route("Summarize the attached note for CNIC 42101-1234567-1", task="summary"))Log the route for every request (without the sensitive content) so you can measure distribution, cost and quality per route.
Gateways such as LiteLLM and several API-management products implement model aliases, fallbacks, budgets and logging for you; the logic above is what you configure in them.
Economics of routing
Model the monthly bill as:
monthly cost ≈ Σ over routes ( requests_route × avg_tokens_route × price_per_token_route ) + fixed local infrastructureIf 70% of requests are simple and move to a local model you already run, and the remaining 30% go to a stronger API, total cost and data exposure both fall, while quality on hard tasks rises. Use your own traffic and current price sheets; figures in vendor examples rarely match your mix.
Worked example: an agency in Lahore serving UK and Gulf clients
- Client-confidential briefs (source metadata = client workspace flagged "NDA"): local only.
- Social captions and hashtag suggestions: local model.
- Quarterly strategy decks for clients who approved cloud processing: stronger API model.
- API outage: local fallback with a banner.
After a month (illustrative): 78% of requests local, 22% API; API spend down versus the previous all-API setup; no confidential brief left the office network; editors rate strategy drafts higher because the strongest model is reserved for them.
Pitfalls
- Running the sensitivity check in the cloud.
- Silent escalation of sensitive data during an outage. Sensitivity rules must beat fallback rules.
- Different system prompts per route causing inconsistent tone; share prompts and test both routes.
- Not logging routes, so you cannot prove where data went.
How to measure success
A route log showing the share of traffic per route, zero sensitive requests on cloud routes (tested with seeded canary data), cost per route, and quality scores per route from your evaluation set.
Key takeaways
- Route per request on sensitivity, task difficulty, cost, latency and availability
- Apply sensitivity rules first; they must override fallbacks
- Detect sensitivity locally with metadata, patterns, a local classifier and a user toggle
- Use validation, self-consistency, retrieval strength and abstention as confidence signals
- Log the route for every request and test with canary data
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Adapt the router to three request types from your organization, seed 20 canary requests with fake IDs, and verify none reach the cloud route.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.