Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APIProduction readiness and capstone · Lesson 19 of 19

Capstone part 2: test, evaluate, deploy and ship a TypeScript client

Article · 22 min · 9 min lecture

Video lecture

Capstone part 2: test, evaluate, deploy and ship a TypeScript client

16 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 16

Capstone part 2

  • Tests with fakes
  • Evals across providers and languages
  • Observability + deployment
  • TypeScript client

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

From working to shippable

Crescent Content API generates content. Now make it trustworthy and easy to use: automated tests without calling real providers, an evaluation harness across providers and languages, observability, a container deployment, and a TypeScript client for the agency's web apps.

Step 1: tests with a fake provider

Test your logic without spending money or depending on provider uptime:

# test_app.py  (pip install pytest httpx)
import json
from fastapi.testclient import TestClient
import app as service
import router
from providers import GenResult, RetryableProviderError

GOOD = json.dumps({"title": "Blue Pottery Mug", "body": "Hand-painted in Multan...", "hashtags": [],
                   "needs_review": False, "review_reasons": []})

class FakeProvider:
    def __init__(self, name, text=GOOD, fail=False): self.name, self.text, self.fail = name, text, fail
    def generate(self, req, model):
        if self.fail:
            raise RetryableProviderError("simulated outage")
        return GenResult(self.text, self.name, model, 100, 50, "end")

client = TestClient(service.app)
BODY = {"tenant_id": "tenant_demo", "task": "product_description", "language": "en",
        "brand_voice": "Warm and modern.", "brief": "Mug, 350ml. Contact ali@example.com"}

def test_happy_path(monkeypatch):
    monkeypatch.setitem(router.PROVIDERS, "anthropic", FakeProvider("anthropic"))
    r = client.post("/v1/generate", json=BODY, headers={"x-api-key": "tenant_demo"})
    assert r.status_code == 200 and r.json()["content"]["title"] == "Blue Pottery Mug"

def test_fallback_on_outage(monkeypatch):
    monkeypatch.setitem(router.PROVIDERS, "anthropic", FakeProvider("anthropic", fail=True))
    monkeypatch.setitem(router.PROVIDERS, "google", FakeProvider("google"))
    r = client.post("/v1/generate", json=BODY, headers={"x-api-key": "tenant_demo"})
    assert r.status_code == 200 and r.json()["provider"] == "google"

def test_invalid_json_is_rejected(monkeypatch):
    monkeypatch.setitem(router.PROVIDERS, "anthropic", FakeProvider("anthropic", text="not json"))
    r = client.post("/v1/generate", json=BODY, headers={"x-api-key": "tenant_demo"})
    assert r.status_code == 502

def test_pii_redacted_before_provider(monkeypatch):
    seen = {}
    class Spy(FakeProvider):
        def generate(self, req, model):
            seen["brief"] = req.messages[0]["text"]
            return super().generate(req, model)
    monkeypatch.setitem(router.PROVIDERS, "anthropic", Spy("anthropic"))
    client.post("/v1/generate", json=BODY, headers={"x-api-key": "tenant_demo"})
    assert "ali@example.com" not in seen["brief"]

(These tests assume router.PROVIDERS is the same registry the router uses; adjust imports to your module layout.)

Step 2: an evaluation harness across providers and languages

Create 30–60 briefs per task across English, Arabic and Urdu, with rubric criteria: brand-voice match, factual fidelity to the brief (no invented claims), length limits, language correctness, needs_review set appropriately (for example a brief claiming "clinically proven" without evidence should be flagged). Score with a calibrated LLM judge plus code checks (length, hashtag count, required #ad when sponsored). Run each route's primary and fallback; record quality, cost per item and latency. Re-run on every prompt, schema or model change; block deploys that regress.

Step 3: observability

  • Structured logs with request_id, tenant, task, provider, model, tokens, cost, latency and outcome (no raw briefs).
  • Metrics: requests, errors by class, fallback activations, breaker state, p50/p95 latency, cost per tenant per day.
  • Traces around provider calls (OpenTelemetry), propagating request_id.
  • Alerts: error rate, fallback rate, spend anomalies, validation failure rate.

Step 4: containerize and deploy

FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
USER 1000
EXPOSE 8080
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8080", "--workers", "2"]

Deploy behind TLS on a container platform in the region your data agreements require; inject provider keys (or workload identity) from a secret manager; run two or more replicas with shared budget state in Redis or your database; set request timeouts above your longest expected generation or use the streaming endpoint.

Step 5: a TypeScript client for web apps

The browser never sees provider keys; it calls your service with the user's session. A small typed client:

// crescentClient.ts
export type Task = 'product_description' | 'social_post' | 'email_draft';
export interface ContentOut { title: string; body: string; hashtags: string[]; needs_review: boolean; review_reasons: string[] }
export interface GenerateOut { content: ContentOut; provider: string; model: string; cost_usd: number; request_id: string }

export async function generate(baseUrl: string, apiKey: string, input: {
  tenant_id: string; task: Task; language: 'en' | 'ar' | 'ur'; brand_voice: string; brief: string;
}, signal?: AbortSignal): Promise<GenerateOut> {
  const res = await fetch(`${baseUrl}/v1/generate`, {
    method: 'POST',
    headers: { 'content-type': 'application/json', 'x-api-key': apiKey },
    body: JSON.stringify(input),
    signal,
  });
  if (res.status === 429) throw new Error('Daily AI budget reached for this workspace.');
  if (res.status === 503) throw new Error('Content generation is temporarily unavailable. Please try again shortly.');
  if (!res.ok) throw new Error(`Generation failed (${res.status}).`);
  return (await res.json()) as GenerateOut;
}

// Streaming variant: consume the SSE relay (lesson 4) from /v1/generate/stream
export function streamDraft(url: string, onDelta: (t: string) => void, onDone: () => void): EventSource {
  const es = new EventSource(url);                      // same-origin, session-authenticated
  es.onmessage = (e) => {
    const data = JSON.parse(e.data);
    if (data.delta) onDelta(data.delta);
    if (data.done || data.error) { es.close(); onDone(); }
  };
  return es;
}

In a server-rendered app, call generate from your server (where the service key lives) rather than the browser. For right-to-left languages, set dir="rtl" on the output container.

Step 6: launch checklist and runbook

Use the production checklist from lesson 17 with evidence, plus a runbook: how to switch a task's primary model (edit routes.yaml, run the eval, deploy), how to force a fallback during an incident, how to raise a tenant's budget, how to rotate keys, and who is on call.

Worked example: the first month

(Illustrative.) Week 1: shadow mode for two clients; eval flags Urdu hashtags as awkward, fixed with a language-specific rule. Week 2: live for five clients; one provider incident triggers the breaker for 12 minutes, fallback handles traffic, clients notice nothing. Week 3: cost report shows social posts cheapest on the fallback at equal quality; routes swapped after re-running the eval. Week 4: all 25 clients live, with per-client AI costs invoiced monthly.

Deliverable

A launch pack: repository with tests, eval results by task, language and provider, dashboards, deployment config, TypeScript client, runbook and checklist with evidence. That pack is the evidence behind this course's badge.

Key takeaways

  • Test service logic with fake providers: happy path, fallback, invalid JSON and redaction.
  • Evaluate every route's primary and fallback across tasks and languages with rubrics and code checks.
  • Observe requests, errors, fallbacks, costs and latency with structured logs, metrics, traces and alerts.
  • Deploy in containers behind TLS with secrets from a manager and shared budget state.
  • Ship a typed TypeScript client that never exposes provider keys, plus a runbook.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. How should unit tests exercise the fallback path?
  2. Where should a browser app get generated content from?
  3. The eval shows the fallback model matches quality for social posts at lower cost. What do you do?

Put it into practice

Add the four tests, run the eval across two providers and three languages, deploy to a staging environment, build a tiny web page using the TypeScript client, and write the runbook.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.