---
title: "Capstone part 1: build a provider-agnostic…"
description: "The brief Crescent Content API (fictional) serves an agency's internal tools and client portals. It generates marketing content (product descriptions…"
url: https://optimizeall.com/learn/ai-platform-apis-integration/capstone-build-content-microservice
updated: 2026-10-05
---

Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Production readiness and capstone · lesson 18 of 19 · 26 min

# Capstone part 1: build a provider-agnostic content-generation microservice

## The brief

**Crescent Content API** (fictional) serves an agency's internal tools and client portals. It generates marketing content (product descriptions, social posts, email drafts) in English, Arabic and Urdu, in each client's brand voice, and must:

- Work with at least two providers (primary + fallback) chosen per task by configuration.
- Return **structured** outputs (title, body, hashtags, compliance flags) validated against schemas.
- Stream long drafts to the UI.
- Track cost per tenant and task; enforce per-tenant daily budgets.
- Redact personal data, keep keys server-side, and log safely.

## Project layout

```text
crescent_content/
  app.py            # FastAPI endpoints
  providers.py      # adapters: Claude, OpenAI, Gemini (lesson 13, extended)
  router.py         # task routes + fallback + circuit breaker
  schemas.py        # Pydantic request/response models
  budget.py         # cost ledger and per-tenant budgets
  redact.py         # PII redaction (lesson 17)
  routes.yaml       # task -> [(provider, model)] configuration
  prices.json       # prices per model (from pricing pages, with effective date)
```

## routes.yaml

```yaml
# Model IDs are configuration. Verify current IDs in each provider's docs before deploying.
product_description:
  - {provider: anthropic, model: claude-sonnet-5}
  - {provider: google, model: gemini-flash-latest}
social_post:
  - {provider: openai, model: gpt-5.5}
  - {provider: anthropic, model: claude-sonnet-5}
email_draft:
  - {provider: anthropic, model: claude-sonnet-5}
  - {provider: openai, model: gpt-5.5}
```

## schemas.py

```python
from typing import Literal
from pydantic import BaseModel, Field

Task = Literal["product_description", "social_post", "email_draft"]
Lang = Literal["en", "ar", "ur"]

class GenerateIn(BaseModel):
    tenant_id: str = Field(min_length=3, max_length=64)
    task: Task
    language: Lang = "en"
    brand_voice: str = Field(max_length=2000, description="Stable brand guidelines (cached)")
    brief: str = Field(max_length=4000, description="What to write about")

class ContentOut(BaseModel):
    title: str = Field(max_length=120)
    body: str = Field(max_length=3000)
    hashtags: list[str] = Field(default_factory=list, max_length=10)
    needs_review: bool = Field(description="True if claims need verification or the brief was unclear")
    review_reasons: list[str] = Field(default_factory=list)

class GenerateOut(BaseModel):
    content: ContentOut
    provider: str
    model: str
    cost_usd: float
    request_id: str

# JSON Schema for providers' strict modes: all properties required, no extra properties
CONTENT_SCHEMA = {
    "type": "object",
    "properties": {
        "title": {"type": "string"}, "body": {"type": "string"},
        "hashtags": {"type": "array", "items": {"type": "string"}},
        "needs_review": {"type": "boolean"},
        "review_reasons": {"type": "array", "items": {"type": "string"}},
    },
    "required": ["title", "body", "hashtags", "needs_review", "review_reasons"],
    "additionalProperties": False,
}
```

## router.py (routing, fallback and a simple circuit breaker)

```python
import time, logging, yaml
from providers import PROVIDERS, GenRequest, GenResult, RetryableProviderError

log = logging.getLogger("router")
ROUTES = yaml.safe_load(open("routes.yaml"))
BREAKER: dict[str, list[float]] = {}          # provider -> recent failure timestamps
OPEN_FOR_S, THRESHOLD, WINDOW_S = 60, 5, 30

def _open(provider: str) -> bool:
    now = time.time()
    fails = [t for t in BREAKER.get(provider, []) if now - t < WINDOW_S]
    BREAKER[provider] = fails
    return len(fails) >= THRESHOLD and now - fails[-1] < OPEN_FOR_S

def generate(task: str, req: GenRequest) -> GenResult:
    last = None
    for step in ROUTES[task]:
        p, model = step["provider"], step["model"]
        if _open(p):
            log.warning("breaker open for %s; skipping", p)
            continue
        try:
            return PROVIDERS[p].generate(req, model)
        except RetryableProviderError as exc:          # timeouts, 429 after retries, 5xx/overload
            BREAKER.setdefault(p, []).append(time.time())
            log.warning("fallback from %s/%s: %s", p, model, exc)
            last = exc
    raise RuntimeError(f"no provider available for {task}") from last
```

`providers.py` is the adapter from lesson 13, extended so each provider wraps its SDK's retryable errors (rate limits after SDK retries, timeouts, connection errors, 5xx) in `RetryableProviderError`, lets non-retryable errors (400s) propagate, and returns normalized usage including cached tokens.

## budget.py

```python
import json, time, threading
from collections import defaultdict

PRICES = json.load(open("prices.json"))          # {"model": {"in": $/1M, "out": $/1M, "cache_read": $/1M}}
DAILY_BUDGET_USD = defaultdict(lambda: 5.0)      # per-tenant default; load real limits from your DB
_spend: dict[tuple[str, str], float] = defaultdict(float)
_lock = threading.Lock()

def today() -> str:
    return time.strftime("%Y-%m-%d")

def check(tenant: str) -> None:
    if _spend[(tenant, today())] >= DAILY_BUDGET_USD[tenant]:
        raise PermissionError("Daily AI budget reached for this tenant")

def record(tenant: str, task: str, model: str, in_tok: int, cached: int, out_tok: int) -> float:
    p = PRICES[model]
    cost = (in_tok * p["in"] + cached * p.get("cache_read", p["in"]) + out_tok * p["out"]) / 1_000_000
    with _lock:
        _spend[(tenant, today())] += cost
    print(json.dumps({"tenant": tenant, "task": task, "model": model, "in": in_tok,
                      "cached": cached, "out": out_tok, "cost": round(cost, 6)}))   # ship to your ledger
    return cost
```

(In production, keep spend in Redis or your database so all replicas share it.)

## app.py

```python
import uuid, logging
from fastapi import FastAPI, HTTPException, Depends, Header
from pydantic import ValidationError
from schemas import GenerateIn, GenerateOut, ContentOut, CONTENT_SCHEMA
from providers import GenRequest
from redact import redact
import router, budget

app = FastAPI(title="Crescent Content API")
log = logging.getLogger("api")

LANG_NAMES = {"en": "English", "ar": "Modern Standard Arabic", "ur": "Urdu"}
TASK_RULES = {
    "product_description": "60-120 words. Concrete benefits. No health, price or performance claims you cannot verify from the brief.",
    "social_post": "Under 60 words plus up to 5 hashtags. Include '#ad' if the brief says the post is sponsored.",
    "email_draft": "Subject-style title under 60 characters; body under 150 words; clear call to action.",
}

def auth(x_api_key: str = Header(...)) -> str:
    # Replace with your real auth (JWT/OAuth); map the caller to a tenant and scopes.
    if not x_api_key.startswith("tenant_"):
        raise HTTPException(401, "invalid credentials")
    return x_api_key

@app.post("/v1/generate", response_model=GenerateOut)
def generate(body: GenerateIn, caller: str = Depends(auth)):
    request_id = str(uuid.uuid4())
    try:
        budget.check(body.tenant_id)
    except PermissionError as e:
        raise HTTPException(429, str(e))
    brief, _ = redact(body.brief)                       # personal data never reaches the provider
    system = (f"You write marketing content for an agency client.\nBrand voice:\n{body.brand_voice}\n"
              f"Rules: {TASK_RULES[body.task]} Write in {LANG_NAMES[body.language]}. "
              "Set needs_review=true with reasons if any claim needs verification.")
    req = GenRequest(system=system, messages=[{"role": "user", "text": brief}], max_tokens=1200,
                     json_schema=CONTENT_SCHEMA, metadata={"tenant": body.tenant_id, "task": body.task})
    try:
        res = router.generate(body.task, req)
    except RuntimeError:
        log.exception("all providers failed request_id=%s", request_id)
        raise HTTPException(503, "content generation temporarily unavailable")
    if res.stop != "end":
        raise HTTPException(502, f"generation incomplete ({res.stop}); try a shorter brief")
    try:
        content = ContentOut.model_validate_json(res.text)
    except ValidationError:
        log.warning("schema validation failed request_id=%s provider=%s", request_id, res.provider)
        raise HTTPException(502, "invalid content returned; please retry")
    cost = budget.record(body.tenant_id, body.task, res.model, res.input_tokens,
                         getattr(res, "cached_tokens", 0), res.output_tokens)
    return GenerateOut(content=content, provider=res.provider, model=res.model,
                       cost_usd=round(cost, 6), request_id=request_id)
```

Note the design: the **brand voice sits in the system prompt** (stable per tenant, so prompt caching can help if you add cache markers in the Claude adapter), the **brief goes in the user turn**, and every response is validated before the client sees it. Add a `/v1/generate/stream` endpoint using the SSE relay from lesson 4 for long drafts, returning a final validated object at the end.

## Run it

```bash
pip install fastapi uvicorn pydantic pyyaml anthropic openai google-genai
export ANTHROPIC_API_KEY=... OPENAI_API_KEY=... GEMINI_API_KEY=...   # or use a secret manager
uvicorn app:app --reload
curl -s localhost:8000/v1/generate -H "x-api-key: tenant_demo" -H "content-type: application/json" \
  -d '{"tenant_id":"tenant_demo","task":"product_description","language":"en",
       "brand_voice":"Warm, modern, proudly handmade in Pakistan.","brief":"Blue pottery mug, 350ml, dishwasher safe"}'
```

Part 2 adds tests, an eval harness, observability, deployment and a TypeScript client.

## Video lecture: Capstone part 1: build a provider-agnostic content-generation microservice

Lecture coming soon · 16 chapters · about 9 minutes. Read the full transcript below.

1. Capstone part 1
2. Why a microservice?
3. The requirements
4. Project layout
5. Schemas
6. Simple example: one request
7. providers.py contract
8. The router
9. Budgets and ledger
10. Business scenario: a busy morning
11. Design choices
12. Common mistakes
13. Why a needs_review flag?
14. Deeper: a morning of traffic (illustrative)
15. Watch me do it: one request through app.py
16. Recap + try this now

## Lecture transcript

### Capstone part 1

Everything in this course comes together now. You'll build Crescent Content API, a provider agnostic microservice that generates marketing content in English, Arabic and Urdu, in each client's brand voice, with structured outputs, fallbacks, streaming, budgets and privacy controls. By the end of this part, you'll have a working service you can call with a single request.

### Why a microservice?

Why build it as a separate service? Because content generation will be used by many tools: a client portal, an internal dashboard, a scheduling app. Centralizing it means one place for keys, routing, budgets, redaction and logs. Think of a hotel's central kitchen serving the restaurant, room service and the banquet hall. Each outlet orders what it needs, but hygiene, stock and costs are managed in one place.

### The requirements

Here's the brief. The service must work with at least two providers per task, chosen by configuration. It returns structured content: a title, a body, hashtags, and a needs review flag with reasons. It streams long drafts. It tracks cost per tenant and task, and enforces daily budgets. And it redacts personal data, keeps keys on the server and logs safely. Every one of those requirements maps to a lesson you've already completed.

### Project layout

The project has a handful of small files. The app file holds the FastAPI endpoints. Providers holds the adapters for Claude, OpenAI and Gemini. Router handles routes, fallback and a simple circuit breaker. Schemas defines the request and response models. Budget keeps the cost ledger and per tenant limits. Redact removes personal data. And two configuration files: routes, mapping each task to a primary and fallback model, and prices, loaded from pricing pages with an effective date.

### Schemas

Start with the schemas. The request carries the tenant, the task, the language, the brand voice and the brief, each with length limits. The content output has a title, body, hashtags, and the needs review flag with reasons, which is the escape hatch that lets the model say, a claim here needs checking. For providers' strict modes, there's also a JSON schema where every property is required and no extra properties are allowed.

### Simple example: one request

A simple example of one request. The client posts: tenant demo, task product description, English, brand voice warm, modern, proudly handmade in Pakistan, and the brief blue pottery mug, three fifty milliliters, dishwasher safe. The service checks the budget, redacts the brief, builds the system prompt with the brand voice and task rules, asks the router for the product description route, validates the JSON, records the cost, and returns the content with the provider, model, cost and a request id.

### providers.py contract

Let's look at the providers file for a moment, because it's where reliability starts. Each adapter wraps its SDK's retryable errors, like rate limits after the SDK's own retries, timeouts, connection errors and server errors, in one common retryable provider error. Non retryable errors, like a bad request, pass through unchanged. It also returns normalized usage, including cached tokens. That small contract is what lets the router make good decisions without knowing anything about Claude, OpenAI or Gemini specifically.

### The router

The router reads routes from configuration. For product descriptions, it tries Claude first and Gemini as the fallback. For social posts, OpenAI first and Claude as the fallback. Before calling a provider, it checks a simple circuit breaker: if a provider has failed five times in thirty seconds, it's skipped for a minute. Only retryable errors, like timeouts, rate limits after retries and server errors, trigger a fallback. Bad requests propagate, because they'd fail everywhere.

### Budgets and ledger

The budget module checks each tenant's spend for the day before calling any provider, and returns a clear four twenty nine style error when the budget is reached. After each successful call, it records input, cached and output tokens with the cost from the prices file, and emits a ledger line. In production, keep that spend in a shared store like Redis or your database, so all replicas see the same numbers.

### Business scenario: a busy morning

Now a realistic business scenario, with illustrative numbers. The agency's portal serves twenty five clients. On a busy morning, a client in Dubai requests Arabic social posts for a Ramadan campaign, a Lahore client requests Urdu product descriptions, and a London client requests English email drafts. Each request goes to its task's primary model. When one provider has a brief incident, the breaker opens and social posts shift to the fallback. Every response is validated, costs are attributed per client, and nothing personal leaves the service unredacted.

### Design choices

Look at the key design choices in the app file. The brand voice sits in the system prompt, because it's stable per tenant, which lets prompt caching help. The brief goes in the user turn. Personal data is redacted before any provider sees it. Incomplete generations and invalid JSON become clear errors instead of broken content. And when every provider fails, the client gets a five oh three with a friendly message, while the logs keep the details.

### Common mistakes

Common mistakes to avoid while building. Putting provider keys in the client portal. Skipping validation because structured outputs usually work. Keeping budget state only in memory on one replica. And hard coding model ids instead of reading them from routes. Each of these shows up in production, usually at the worst time.

### Why a needs_review flag?

Why does the service return a needs review flag instead of just refusing risky content? Because in marketing, many briefs are fine with small edits, and a human reviewer can fix them quickly. The flag, with reasons, turns the model into a helpful first drafter that also highlights risk, like an unsupported health claim or a missing sponsorship label. Your portal can show flagged drafts with a yellow banner, so reviewers know exactly where to look.

### Deeper: a morning of traffic (illustrative)

Let's deepen the Crescent Content API with a realistic morning, illustrative numbers. Between nine and eleven, the portal handles a few hundred requests from fourteen clients. A Lahore home goods brand generates forty Urdu product descriptions; the router uses Claude, and six are flagged for review because the briefs mentioned health benefits. A Dubai café chain generates Arabic social posts for a Ramadan campaign; the sponsored ones include the ad tag automatically because of the task rules. A London fitness studio's budget hits its daily limit at ten thirty, and its marketer sees a clear message instead of an error page. Every request has a request id, a provider, a model and a cost in the ledger.

### Watch me do it: one request through app.py

Watch me do it. Let's trace one request through app dot py. The auth dependency checks the API key header and maps it to a caller. Budget check raises if the tenant's spend today has reached its limit, which becomes a four twenty nine. Redact removes personal data from the brief. The system prompt combines the brand voice, the task rules and the language name. The request object carries the system prompt, one user message, a twelve hundred token limit, the strict content schema and metadata. Router generate reads the routes for this task, skips any provider whose breaker is open, and returns the first success. If every provider fails, the endpoint returns a five oh three. If the stop value isn't end, it returns a five oh two. Content out model validate json checks the JSON against the schema. Budget record writes the ledger line and returns the cost. The response includes the content, provider, model, cost and request id.

### Recap + try this now

Quick recap. You've built a multi provider content service with configuration based routing, fallbacks and a circuit breaker, strict structured outputs validated with Pydantic, budgets and a cost ledger, redaction and server side keys. Try this now: build it with at least two providers, set up the routes file, and generate one piece of content for each task in two languages, recording the provider and cost for each. In part two, you'll test, evaluate, observe, deploy and add a TypeScript client.

## Key takeaways

- The capstone service routes each task to a primary and fallback provider from configuration.
- Structured outputs are requested via one strict JSON Schema and validated with Pydantic before returning.
- Budgets, cost recording, redaction and backend-only keys are built in from the start.
- Stable brand voice goes in the system prompt; the variable brief goes in the user turn.
- Failures map to clear HTTP errors: 429 for budgets, 502 for invalid output, 503 when all providers fail.

## Try it

Build the service with at least two providers, configure routes.yaml, and generate one piece of content per task in two languages. Record cost and provider for each.

- [Previous: Security, data privacy and the production checklist](https://optimizeall.com/learn/ai-platform-apis-integration/security-privacy-production-checklist)
- [Next: Capstone part 2: test, evaluate, deploy and ship a TypeScript client](https://optimizeall.com/learn/ai-platform-apis-integration/capstone-test-deploy-typescript-client)
- [All lessons of Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API](https://optimizeall.com/learn/ai-platform-apis-integration)
