Skip to content

Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Multi-provider architecture · lesson 13 of 19 · 17 min

Multi-provider abstraction and fallbacks

Why abstract

Models improve and prices change every few months; outages happen; one provider is best for Urdu copy, another for extraction. A thin abstraction layer lets you switch models per task by configuration, fail over during incidents, and compare providers on the same eval, without rewriting business logic.

Build, buy or both

  • Your own thin adapter: a small interface (generate, stream, generate_structured, embed) with one implementation per provider using the official SDKs. Maximum control, minimal dependencies, full access to provider-specific features.
  • Open-source libraries and gateways (for example LiteLLM, or framework abstractions in LangChain, LlamaIndex, Pydantic AI and the Vercel AI SDK): many providers behind one API, often with routing, retries and cost tracking. Faster to start; check how well they support the newest features (caching controls, structured outputs, thinking, server tools) and how quickly they update.
  • Managed AI gateways from cloud and CDN providers or specialist vendors: central keys, logging, rate limits and fallbacks as a service.
  • OpenAI-compatible endpoints: many providers (including Google's Gemini API compatibility layer and most open-model hosts) accept Chat Completions-shaped requests. Useful for quick switching, but compatibility layers usually lag or omit provider-specific features, and subtle behavioral differences remain. For Claude, prefer the official Anthropic SDK or its cloud-platform clients rather than compatibility shims.

A common pattern: your own adapter interface, implemented with official SDKs for your top two providers, plus an OpenAI-compatible implementation for long-tail open models.

The adapter contract

Keep the interface small and honest:

from dataclasses import dataclass, field
from typing import Iterator, Protocol

@dataclass
class GenRequest:
    system: str
    messages: list[dict]                 # [{"role": "user"|"assistant", "text": "..."}]
    max_tokens: int = 1024
    json_schema: dict | None = None      # structured output if set
    metadata: dict = field(default_factory=dict)   # feature, tenant for cost tracking

@dataclass
class GenResult:
    text: str
    provider: str
    model: str
    input_tokens: int
    output_tokens: int
    stop: str                            # "end" | "length" | "refusal" | "error"

class Provider(Protocol):
    name: str
    def generate(self, req: GenRequest, model: str) -> GenResult: ...
    def stream(self, req: GenRequest, model: str) -> Iterator[str]: ...

Normalize stop reasons and usage so upper layers can make decisions (retry, fall back, bill) without knowing the provider.

Implementations (abbreviated)

import os, anthropic
from openai import OpenAI
from google import genai
from google.genai import types

class ClaudeProvider:
    name = "anthropic"
    def __init__(self): self.c = anthropic.Anthropic()
    def generate(self, req, model):
        kw = {}
        if req.json_schema:
            kw["output_config"] = {"format": {"type": "json_schema", "schema": req.json_schema}}
        r = self.c.messages.create(model=model, max_tokens=req.max_tokens, system=req.system,
                                   messages=[{"role": m["role"], "content": m["text"]} for m in req.messages], **kw)
        stop = {"end_turn": "end", "max_tokens": "length", "refusal": "refusal"}.get(r.stop_reason, r.stop_reason)
        return GenResult("".join(b.text for b in r.content if b.type == "text"), self.name, model,
                         r.usage.input_tokens, r.usage.output_tokens, stop)

class OpenAIProvider:
    name = "openai"
    def __init__(self): self.c = OpenAI()
    def generate(self, req, model):
        kw = {}
        if req.json_schema:
            kw["text"] = {"format": {"type": "json_schema", "name": "out", "schema": req.json_schema, "strict": True}}
        r = self.c.responses.create(model=model, instructions=req.system, max_output_tokens=req.max_tokens,
                                    input=[{"role": m["role"], "content": m["text"]} for m in req.messages],
                                    store=False, **kw)
        stop = "length" if getattr(r, "incomplete_details", None) else "end"
        return GenResult(r.output_text, self.name, model, r.usage.input_tokens, r.usage.output_tokens, stop)

class GeminiProvider:
    name = "google"
    def __init__(self): self.c = genai.Client()
    def generate(self, req, model):
        cfg = types.GenerateContentConfig(system_instruction=req.system, max_output_tokens=req.max_tokens)
        if req.json_schema:
            cfg.response_mime_type, cfg.response_json_schema = "application/json", req.json_schema
        contents = [types.Content(role="model" if m["role"] == "assistant" else "user",
                                  parts=[types.Part.from_text(text=m["text"])]) for m in req.messages]
        r = self.c.models.generate_content(model=model, contents=contents, config=cfg)
        u = r.usage_metadata
        return GenResult(r.text or "", self.name, model, u.prompt_token_count or 0,
                         u.candidates_token_count or 0, "end")

Structured-output schemas must meet each provider's strict-mode rules (for example additionalProperties: false and required fields for OpenAI strict mode). Gemini finish reasons and safety blocks should also be mapped; they are omitted here for brevity.

Fallback chains

import logging
log = logging.getLogger("router")
PROVIDERS = {"anthropic": ClaudeProvider(), "openai": OpenAIProvider(), "google": GeminiProvider()}
ROUTES = {   # configuration, not code: loaded from YAML/DB in production
  "product_copy": [("anthropic", os.environ.get("COPY_PRIMARY", "claude-sonnet-5")),
                   ("google", os.environ.get("COPY_FALLBACK", "gemini-flash-latest"))],
}

def generate(task: str, req: GenRequest) -> GenResult:
    last_exc = None
    for provider_name, model in ROUTES[task]:
        try:
            res = PROVIDERS[provider_name].generate(req, model)
            if res.stop == "refusal":
                log.info("refusal from %s/%s; not retrying elsewhere automatically", provider_name, model)
            return res
        except Exception as exc:          # narrow to retryable provider errors in real code
            log.warning("fallback: %s/%s failed: %s", provider_name, model, type(exc).__name__)
            last_exc = exc
    raise RuntimeError(f"All providers failed for {task}") from last_exc

Decide deliberately what triggers a fallback: provider outages, overloads and timeouts usually should; bad requests should not (they would fail everywhere); refusals need a product decision (some providers now offer documented server-side fallback mechanisms for refusals; check before building your own).

Keep fallbacks honest

  • Evaluate every model in a chain on the same eval; a fallback that produces worse or off-brand output needs guardrails or a narrower role.
  • Prompt portability: prompts tuned for one model may underperform on another; keep per-model prompt variants where it matters.
  • Data terms: failing over to a provider not covered by your data agreements can breach commitments. Only include providers approved for that data class.
  • Caching: fallbacks forfeit cache hits on the primary; expect higher cost during incidents.

Worked example: a Riyadh news-summary service

A media startup summarizes Arabic and English news for subscribers. Primary: provider A's mid-tier model (best Arabic quality on their eval). Fallback: provider B's model with an Arabic-specific prompt variant, approved under the same data agreement. During a primary-provider incident (illustrative), the circuit breaker opened, traffic moved to the fallback within seconds, and subscribers noticed only slightly different phrasing. Weekly evals run on both so the fallback stays fit for purpose.

Pitfalls

  • Lowest-common-denominator abstractions that block useful provider features.
  • Untested fallbacks discovered to be broken during an outage.
  • Falling back on errors that will fail everywhere (400s).
  • Ignoring data-agreement boundaries when routing.

Measuring success

Fallback activation rate and success rate, quality delta between primary and fallback on your eval, time to switch a task's primary model (config + eval), and incident impact (failed requests during outages).

Video lecture: Multi-provider abstraction and fallbacks

Lecture coming soon · 16 chapters · about 9 minutes. Read the full transcript below.

  1. Abstraction and fallbacks
  2. Why it matters
  3. Build, buy or both
  4. A balanced pattern
  5. The adapter contract
  6. Simple example: tone rewrite
  7. Three implementations
  8. Fallback rules
  9. Business example: Riyadh news summaries
  10. Honest fallbacks
  11. Common mistakes
  12. Library or gateway?
  13. Monthly fallback game day
  14. Deeper: the 6 a.m. incident (illustrative)
  15. Watch me do it: adapter + router
  16. Recap + try this now

Lecture transcript

Abstraction and fallbacks

What happens to your product when your AI provider has an outage, raises prices, or releases a model that's worse at your language? If the answer is, we'd have to rewrite a lot of code, this lesson is for you. You'll learn how to build a thin abstraction layer across providers, when to use libraries or gateways instead, and how to design fallbacks that actually work when you need them.

Why it matters

Why does this matter? Models improve and prices change every few months, outages happen, and different providers are best at different tasks. Think of a power strip with universal sockets. Your devices don't care which brand of plug is in the wall, because the adapter handles it. An abstraction layer is that power strip for AI: your business logic plugs into one interface, and providers plug in behind it.

Build, buy or both

You have four options. Build your own thin adapter with the official SDKs, which gives maximum control and full access to each provider's features. Use an open source library or gateway, like LiteLLM or the abstractions in popular frameworks, which is faster to start but may lag on the newest features. Use a managed AI gateway for central keys, logging and fallbacks as a service. Or use OpenAI compatible endpoints, which many providers offer. They're handy for quick switching, but often lag or omit provider specific features.

A balanced pattern

A common, balanced pattern: define your own small interface, implement it with official SDKs for your top two providers, and add an OpenAI compatible implementation for long tail open models. And for Claude specifically, prefer the official Anthropic SDK or its cloud platform clients rather than compatibility shims, so features like caching controls, structured outputs and thinking work properly.

The adapter contract

What goes in the interface? Keep it small and honest. A request with a system prompt, messages, a max tokens value, an optional JSON schema for structured output, and metadata like feature and tenant for cost tracking. A result with the text, the provider and model used, input and output tokens, and a normalized stop reason: end, length, refusal or error. Normalizing stop reasons and usage is what lets your upper layers decide on retries, fallbacks and billing without caring which provider answered.

Simple example: tone rewrite

A simple example. A small startup has one feature: rewrite a paragraph in a friendlier tone. Their code calls one generate function with the task name. A configuration file says: use provider A's model, and if it fails, use provider B's. When they later find provider B is just as good and cheaper, they swap the order in the config and rerun their eval. No code changes, no deployment of new logic, just configuration and evidence.

Three implementations

The lesson shows implementations for Claude, OpenAI and Gemini. Each converts the common request into the provider's format, adds structured output settings when a schema is given, and converts the response back, mapping stop reasons and usage. Watch the details: Claude needs max tokens, OpenAI's Responses API takes instructions and input, with store set to false here, and Gemini uses model instead of assistant for earlier turns. Schemas must also satisfy each provider's strict mode rules.

Fallback rules

Now fallbacks. Define routes per task in configuration: a primary model and one or more fallbacks. The router tries each in order and logs every fallback. But decide carefully what triggers a fallback. Outages, overloads and timeouts should. Bad requests shouldn't, because they'd fail everywhere. And refusals need a product decision rather than an automatic retry elsewhere. Some providers now offer documented server side fallback options for refusals, so check before building your own.

Business example: Riyadh news summaries

Now a realistic business example, with illustrative numbers. A Riyadh media startup summarizes Arabic and English news for subscribers. Its primary is provider A's mid tier model, which won their Arabic eval. The fallback is provider B's model with an Arabic specific prompt variant, approved under the same data agreement. During a primary provider incident, the circuit breaker opened, traffic moved to the fallback within seconds, and subscribers noticed only slightly different phrasing. They run weekly evals on both, so the fallback stays fit for purpose.

Honest fallbacks

Keep your fallbacks honest. Evaluate every model in the chain on the same eval. Maintain per model prompt variants where quality differs. Only route data to providers covered by your data agreements, because failing over to an unapproved provider during an outage can breach commitments. And expect higher cost during incidents, because you lose cache hits on the primary.

Common mistakes

Common mistakes. Lowest common denominator abstractions that block useful provider features. Untested fallbacks that turn out to be broken during an outage. Falling back on errors that will fail everywhere. And ignoring data agreement boundaries when routing. Any one of these can turn a well intentioned abstraction into a source of outages, poor quality or compliance problems.

Library or gateway?

A question architects ask: where should the abstraction live, in every service or in one gateway? For a small team, a shared library is simplest. As you grow, a central internal gateway service often wins: one place for keys, routing, budgets, logging and fallbacks, used by every product team over a simple internal API. Either way, keep the interface small, version it, and make provider specific features available when a team genuinely needs them.

Monthly fallback game day

A quick word on testing fallbacks, because untested fallbacks are the most common surprise during real outages. Schedule a monthly game day. In staging, block the primary provider with a firewall rule or an invalid endpoint, and run your eval suite through the router. Check three things: that the fallback activates within seconds, that quality stays inside your agreed range, and that logs and cost reports clearly show which requests used the fallback. Then restore the primary and check traffic returns.

Deeper: the 6 a.m. incident (illustrative)

Let's deepen the Riyadh news summary example with illustrative numbers. The service sends morning summaries to subscribers at six a.m. During one primary provider incident, requests started timing out at five fifty. The breaker opened after five failures within thirty seconds, and the router sent traffic to the fallback with its Arabic prompt variant. The six a.m. send went out on time. Weekly evals had shown the fallback scored a few points lower on Arabic fluency, so the editor on duty spot checked the summaries that morning and found them acceptable. When the primary recovered, traffic returned automatically, and the incident report was two paragraphs long instead of an apology email to subscribers.

Watch me do it: adapter + router

Watch me do it. Let's walk through the adapter and router. GenRequest carries the system prompt, messages, max tokens, an optional JSON schema and metadata. GenResult carries text, provider, model, tokens and a normalized stop value. The Claude provider builds messages create arguments, adds output config when there's a schema, and maps end turn to end, max tokens to length and refusal to refusal. The OpenAI provider uses instructions, input, max output tokens and store false, adds a text format when there's a schema, and marks length when the response is incomplete. The Gemini provider builds a config, renames assistant to model, and reads usage metadata. The router has routes per task: product copy tries Claude, then Gemini. Generate loops through the route, returns the first success, logs any fallback, and raises if all fail. I point the primary at a broken endpoint and rerun: the log says fallback, and the result comes from Gemini.

Recap + try this now

Quick recap. A thin adapter makes models swappable by configuration. Normalize usage and stop reasons, route per task with fallbacks for outages, and keep fallbacks evaluated, prompt tuned and within your data agreements. Try this now: implement the adapter for two providers, configure a primary and fallback for one task, run your eval on both, and simulate an outage by pointing the primary at an invalid endpoint to confirm the fallback activates.

Key takeaways

  • A thin adapter layer lets you switch models by configuration, fail over and compare providers fairly.
  • Choose between your own adapter, open-source libraries or gateways, managed gateways and OpenAI-compatible endpoints based on feature needs.
  • Normalize text, usage and stop reasons so upper layers can decide on retries, fallbacks and billing.
  • Fall back on outages and timeouts, not on bad requests; treat refusals as a product decision.
  • Evaluate every model in a chain and respect data-agreement boundaries.

Try it

Implement the adapter for two providers, configure a primary and fallback for one task, run your eval on both, and simulate an outage to confirm the fallback activates.