Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APIProduction readiness and capstone · Lesson 18 of 19

Capstone part 1: build a provider-agnostic content-generation microservice

Article · 26 min · 9 min lecture

Video lecture

Capstone part 1: build a provider-agnostic content-generation microservice

16 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 16

Capstone part 1

  • Crescent Content API
  • Multi-provider, structured, budgeted
  • EN, AR, UR content

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The brief

Crescent Content API (fictional) serves an agency's internal tools and client portals. It generates marketing content (product descriptions, social posts, email drafts) in English, Arabic and Urdu, in each client's brand voice, and must:

  • Work with at least two providers (primary + fallback) chosen per task by configuration.
  • Return structured outputs (title, body, hashtags, compliance flags) validated against schemas.
  • Stream long drafts to the UI.
  • Track cost per tenant and task; enforce per-tenant daily budgets.
  • Redact personal data, keep keys server-side, and log safely.

Project layout

crescent_content/
  app.py            # FastAPI endpoints
  providers.py      # adapters: Claude, OpenAI, Gemini (lesson 13, extended)
  router.py         # task routes + fallback + circuit breaker
  schemas.py        # Pydantic request/response models
  budget.py         # cost ledger and per-tenant budgets
  redact.py         # PII redaction (lesson 17)
  routes.yaml       # task -> [(provider, model)] configuration
  prices.json       # prices per model (from pricing pages, with effective date)

routes.yaml

# Model IDs are configuration. Verify current IDs in each provider's docs before deploying.
product_description:
  - {provider: anthropic, model: claude-sonnet-5}
  - {provider: google, model: gemini-flash-latest}
social_post:
  - {provider: openai, model: gpt-5.5}
  - {provider: anthropic, model: claude-sonnet-5}
email_draft:
  - {provider: anthropic, model: claude-sonnet-5}
  - {provider: openai, model: gpt-5.5}

schemas.py

from typing import Literal
from pydantic import BaseModel, Field

Task = Literal["product_description", "social_post", "email_draft"]
Lang = Literal["en", "ar", "ur"]

class GenerateIn(BaseModel):
    tenant_id: str = Field(min_length=3, max_length=64)
    task: Task
    language: Lang = "en"
    brand_voice: str = Field(max_length=2000, description="Stable brand guidelines (cached)")
    brief: str = Field(max_length=4000, description="What to write about")

class ContentOut(BaseModel):
    title: str = Field(max_length=120)
    body: str = Field(max_length=3000)
    hashtags: list[str] = Field(default_factory=list, max_length=10)
    needs_review: bool = Field(description="True if claims need verification or the brief was unclear")
    review_reasons: list[str] = Field(default_factory=list)

class GenerateOut(BaseModel):
    content: ContentOut
    provider: str
    model: str
    cost_usd: float
    request_id: str

# JSON Schema for providers' strict modes: all properties required, no extra properties
CONTENT_SCHEMA = {
    "type": "object",
    "properties": {
        "title": {"type": "string"}, "body": {"type": "string"},
        "hashtags": {"type": "array", "items": {"type": "string"}},
        "needs_review": {"type": "boolean"},
        "review_reasons": {"type": "array", "items": {"type": "string"}},
    },
    "required": ["title", "body", "hashtags", "needs_review", "review_reasons"],
    "additionalProperties": False,
}

router.py (routing, fallback and a simple circuit breaker)

import time, logging, yaml
from providers import PROVIDERS, GenRequest, GenResult, RetryableProviderError

log = logging.getLogger("router")
ROUTES = yaml.safe_load(open("routes.yaml"))
BREAKER: dict[str, list[float]] = {}          # provider -> recent failure timestamps
OPEN_FOR_S, THRESHOLD, WINDOW_S = 60, 5, 30

def _open(provider: str) -> bool:
    now = time.time()
    fails = [t for t in BREAKER.get(provider, []) if now - t < WINDOW_S]
    BREAKER[provider] = fails
    return len(fails) >= THRESHOLD and now - fails[-1] < OPEN_FOR_S

def generate(task: str, req: GenRequest) -> GenResult:
    last = None
    for step in ROUTES[task]:
        p, model = step["provider"], step["model"]
        if _open(p):
            log.warning("breaker open for %s; skipping", p)
            continue
        try:
            return PROVIDERS[p].generate(req, model)
        except RetryableProviderError as exc:          # timeouts, 429 after retries, 5xx/overload
            BREAKER.setdefault(p, []).append(time.time())
            log.warning("fallback from %s/%s: %s", p, model, exc)
            last = exc
    raise RuntimeError(f"no provider available for {task}") from last

providers.py is the adapter from lesson 13, extended so each provider wraps its SDK's retryable errors (rate limits after SDK retries, timeouts, connection errors, 5xx) in RetryableProviderError, lets non-retryable errors (400s) propagate, and returns normalized usage including cached tokens.

budget.py

import json, time, threading
from collections import defaultdict

PRICES = json.load(open("prices.json"))          # {"model": {"in": $/1M, "out": $/1M, "cache_read": $/1M}}
DAILY_BUDGET_USD = defaultdict(lambda: 5.0)      # per-tenant default; load real limits from your DB
_spend: dict[tuple[str, str], float] = defaultdict(float)
_lock = threading.Lock()

def today() -> str:
    return time.strftime("%Y-%m-%d")

def check(tenant: str) -> None:
    if _spend[(tenant, today())] >= DAILY_BUDGET_USD[tenant]:
        raise PermissionError("Daily AI budget reached for this tenant")

def record(tenant: str, task: str, model: str, in_tok: int, cached: int, out_tok: int) -> float:
    p = PRICES[model]
    cost = (in_tok * p["in"] + cached * p.get("cache_read", p["in"]) + out_tok * p["out"]) / 1_000_000
    with _lock:
        _spend[(tenant, today())] += cost
    print(json.dumps({"tenant": tenant, "task": task, "model": model, "in": in_tok,
                      "cached": cached, "out": out_tok, "cost": round(cost, 6)}))   # ship to your ledger
    return cost

(In production, keep spend in Redis or your database so all replicas share it.)

app.py

import uuid, logging
from fastapi import FastAPI, HTTPException, Depends, Header
from pydantic import ValidationError
from schemas import GenerateIn, GenerateOut, ContentOut, CONTENT_SCHEMA
from providers import GenRequest
from redact import redact
import router, budget

app = FastAPI(title="Crescent Content API")
log = logging.getLogger("api")

LANG_NAMES = {"en": "English", "ar": "Modern Standard Arabic", "ur": "Urdu"}
TASK_RULES = {
    "product_description": "60-120 words. Concrete benefits. No health, price or performance claims you cannot verify from the brief.",
    "social_post": "Under 60 words plus up to 5 hashtags. Include '#ad' if the brief says the post is sponsored.",
    "email_draft": "Subject-style title under 60 characters; body under 150 words; clear call to action.",
}

def auth(x_api_key: str = Header(...)) -> str:
    # Replace with your real auth (JWT/OAuth); map the caller to a tenant and scopes.
    if not x_api_key.startswith("tenant_"):
        raise HTTPException(401, "invalid credentials")
    return x_api_key

@app.post("/v1/generate", response_model=GenerateOut)
def generate(body: GenerateIn, caller: str = Depends(auth)):
    request_id = str(uuid.uuid4())
    try:
        budget.check(body.tenant_id)
    except PermissionError as e:
        raise HTTPException(429, str(e))
    brief, _ = redact(body.brief)                       # personal data never reaches the provider
    system = (f"You write marketing content for an agency client.\nBrand voice:\n{body.brand_voice}\n"
              f"Rules: {TASK_RULES[body.task]} Write in {LANG_NAMES[body.language]}. "
              "Set needs_review=true with reasons if any claim needs verification.")
    req = GenRequest(system=system, messages=[{"role": "user", "text": brief}], max_tokens=1200,
                     json_schema=CONTENT_SCHEMA, metadata={"tenant": body.tenant_id, "task": body.task})
    try:
        res = router.generate(body.task, req)
    except RuntimeError:
        log.exception("all providers failed request_id=%s", request_id)
        raise HTTPException(503, "content generation temporarily unavailable")
    if res.stop != "end":
        raise HTTPException(502, f"generation incomplete ({res.stop}); try a shorter brief")
    try:
        content = ContentOut.model_validate_json(res.text)
    except ValidationError:
        log.warning("schema validation failed request_id=%s provider=%s", request_id, res.provider)
        raise HTTPException(502, "invalid content returned; please retry")
    cost = budget.record(body.tenant_id, body.task, res.model, res.input_tokens,
                         getattr(res, "cached_tokens", 0), res.output_tokens)
    return GenerateOut(content=content, provider=res.provider, model=res.model,
                       cost_usd=round(cost, 6), request_id=request_id)

Note the design: the brand voice sits in the system prompt (stable per tenant, so prompt caching can help if you add cache markers in the Claude adapter), the brief goes in the user turn, and every response is validated before the client sees it. Add a /v1/generate/stream endpoint using the SSE relay from lesson 4 for long drafts, returning a final validated object at the end.

Run it

pip install fastapi uvicorn pydantic pyyaml anthropic openai google-genai
export ANTHROPIC_API_KEY=... OPENAI_API_KEY=... GEMINI_API_KEY=...   # or use a secret manager
uvicorn app:app --reload
curl -s localhost:8000/v1/generate -H "x-api-key: tenant_demo" -H "content-type: application/json" \
  -d '{"tenant_id":"tenant_demo","task":"product_description","language":"en",
       "brand_voice":"Warm, modern, proudly handmade in Pakistan.","brief":"Blue pottery mug, 350ml, dishwasher safe"}'

Part 2 adds tests, an eval harness, observability, deployment and a TypeScript client.

Key takeaways

  • The capstone service routes each task to a primary and fallback provider from configuration.
  • Structured outputs are requested via one strict JSON Schema and validated with Pydantic before returning.
  • Budgets, cost recording, redaction and backend-only keys are built in from the start.
  • Stable brand voice goes in the system prompt; the variable brief goes in the user turn.
  • Failures map to clear HTTP errors: 429 for budgets, 502 for invalid output, 503 when all providers fail.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why is the brand voice placed in the system prompt and the brief in the user turn?
  2. A provider returns JSON that fails Pydantic validation. What does the service do?
  3. What should happen when a tenant exceeds its daily budget?

Put it into practice

Build the service with at least two providers, configure routes.yaml, and generate one piece of content per task in two languages. Record cost and provider for each.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.