---
title: "Capstone part 2: test, evaluate, deploy and ship a…"
description: "From working to shippable Crescent Content API generates content. Now make it trustworthy and easy to use: automated tests without calling real…"
url: https://optimizeall.com/learn/ai-platform-apis-integration/capstone-test-deploy-typescript-client
updated: 2026-10-05
---

Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Production readiness and capstone · lesson 19 of 19 · 22 min

# Capstone part 2: test, evaluate, deploy and ship a TypeScript client

## From working to shippable

Crescent Content API generates content. Now make it trustworthy and easy to use: automated tests without calling real providers, an evaluation harness across providers and languages, observability, a container deployment, and a TypeScript client for the agency's web apps.

## Step 1: tests with a fake provider

Test your logic without spending money or depending on provider uptime:

```python
# test_app.py  (pip install pytest httpx)
import json
from fastapi.testclient import TestClient
import app as service
import router
from providers import GenResult, RetryableProviderError

GOOD = json.dumps({"title": "Blue Pottery Mug", "body": "Hand-painted in Multan...", "hashtags": [],
                   "needs_review": False, "review_reasons": []})

class FakeProvider:
    def __init__(self, name, text=GOOD, fail=False): self.name, self.text, self.fail = name, text, fail
    def generate(self, req, model):
        if self.fail:
            raise RetryableProviderError("simulated outage")
        return GenResult(self.text, self.name, model, 100, 50, "end")

client = TestClient(service.app)
BODY = {"tenant_id": "tenant_demo", "task": "product_description", "language": "en",
        "brand_voice": "Warm and modern.", "brief": "Mug, 350ml. Contact ali@example.com"}

def test_happy_path(monkeypatch):
    monkeypatch.setitem(router.PROVIDERS, "anthropic", FakeProvider("anthropic"))
    r = client.post("/v1/generate", json=BODY, headers={"x-api-key": "tenant_demo"})
    assert r.status_code == 200 and r.json()["content"]["title"] == "Blue Pottery Mug"

def test_fallback_on_outage(monkeypatch):
    monkeypatch.setitem(router.PROVIDERS, "anthropic", FakeProvider("anthropic", fail=True))
    monkeypatch.setitem(router.PROVIDERS, "google", FakeProvider("google"))
    r = client.post("/v1/generate", json=BODY, headers={"x-api-key": "tenant_demo"})
    assert r.status_code == 200 and r.json()["provider"] == "google"

def test_invalid_json_is_rejected(monkeypatch):
    monkeypatch.setitem(router.PROVIDERS, "anthropic", FakeProvider("anthropic", text="not json"))
    r = client.post("/v1/generate", json=BODY, headers={"x-api-key": "tenant_demo"})
    assert r.status_code == 502

def test_pii_redacted_before_provider(monkeypatch):
    seen = {}
    class Spy(FakeProvider):
        def generate(self, req, model):
            seen["brief"] = req.messages[0]["text"]
            return super().generate(req, model)
    monkeypatch.setitem(router.PROVIDERS, "anthropic", Spy("anthropic"))
    client.post("/v1/generate", json=BODY, headers={"x-api-key": "tenant_demo"})
    assert "ali@example.com" not in seen["brief"]
```

(These tests assume `router.PROVIDERS` is the same registry the router uses; adjust imports to your module layout.)

## Step 2: an evaluation harness across providers and languages

Create 30–60 briefs per task across English, Arabic and Urdu, with rubric criteria: brand-voice match, factual fidelity to the brief (no invented claims), length limits, language correctness, `needs_review` set appropriately (for example a brief claiming "clinically proven" without evidence should be flagged). Score with a calibrated LLM judge plus code checks (length, hashtag count, required `#ad` when sponsored). Run each route's primary and fallback; record quality, cost per item and latency. Re-run on every prompt, schema or model change; block deploys that regress.

## Step 3: observability

- Structured logs with `request_id`, tenant, task, provider, model, tokens, cost, latency and outcome (no raw briefs).
- Metrics: requests, errors by class, fallback activations, breaker state, p50/p95 latency, cost per tenant per day.
- Traces around provider calls (OpenTelemetry), propagating `request_id`.
- Alerts: error rate, fallback rate, spend anomalies, validation failure rate.

## Step 4: containerize and deploy

```dockerfile
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
USER 1000
EXPOSE 8080
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8080", "--workers", "2"]
```

Deploy behind TLS on a container platform in the region your data agreements require; inject provider keys (or workload identity) from a secret manager; run two or more replicas with shared budget state in Redis or your database; set request timeouts above your longest expected generation or use the streaming endpoint.

## Step 5: a TypeScript client for web apps

The browser never sees provider keys; it calls your service with the user's session. A small typed client:

```typescript
// crescentClient.ts
export type Task = 'product_description' | 'social_post' | 'email_draft';
export interface ContentOut { title: string; body: string; hashtags: string[]; needs_review: boolean; review_reasons: string[] }
export interface GenerateOut { content: ContentOut; provider: string; model: string; cost_usd: number; request_id: string }

export async function generate(baseUrl: string, apiKey: string, input: {
  tenant_id: string; task: Task; language: 'en' | 'ar' | 'ur'; brand_voice: string; brief: string;
}, signal?: AbortSignal): Promise<GenerateOut> {
  const res = await fetch(`${baseUrl}/v1/generate`, {
    method: 'POST',
    headers: { 'content-type': 'application/json', 'x-api-key': apiKey },
    body: JSON.stringify(input),
    signal,
  });
  if (res.status === 429) throw new Error('Daily AI budget reached for this workspace.');
  if (res.status === 503) throw new Error('Content generation is temporarily unavailable. Please try again shortly.');
  if (!res.ok) throw new Error(`Generation failed (${res.status}).`);
  return (await res.json()) as GenerateOut;
}

// Streaming variant: consume the SSE relay (lesson 4) from /v1/generate/stream
export function streamDraft(url: string, onDelta: (t: string) => void, onDone: () => void): EventSource {
  const es = new EventSource(url);                      // same-origin, session-authenticated
  es.onmessage = (e) => {
    const data = JSON.parse(e.data);
    if (data.delta) onDelta(data.delta);
    if (data.done || data.error) { es.close(); onDone(); }
  };
  return es;
}
```

In a server-rendered app, call `generate` from your server (where the service key lives) rather than the browser. For right-to-left languages, set `dir="rtl"` on the output container.

## Step 6: launch checklist and runbook

Use the production checklist from lesson 17 with evidence, plus a runbook: how to switch a task's primary model (edit `routes.yaml`, run the eval, deploy), how to force a fallback during an incident, how to raise a tenant's budget, how to rotate keys, and who is on call.

## Worked example: the first month

(Illustrative.) Week 1: shadow mode for two clients; eval flags Urdu hashtags as awkward, fixed with a language-specific rule. Week 2: live for five clients; one provider incident triggers the breaker for 12 minutes, fallback handles traffic, clients notice nothing. Week 3: cost report shows social posts cheapest on the fallback at equal quality; routes swapped after re-running the eval. Week 4: all 25 clients live, with per-client AI costs invoiced monthly.

## Deliverable

A launch pack: repository with tests, eval results by task, language and provider, dashboards, deployment config, TypeScript client, runbook and checklist with evidence. That pack is the evidence behind this course's badge.

## Video lecture: Capstone part 2: test, evaluate, deploy and ship a TypeScript client

Lecture coming soon · 16 chapters · about 9 minutes. Read the full transcript below.

1. Capstone part 2
2. Why fakes?
3. Four tests
4. Eval harness
5. Simple eval example
6. Observability
7. Reading eval results
8. Deployment
9. TypeScript client
10. Business story: the first month
11. Runbook essentials
12. Common mistakes
13. Adding a new task
14. Deeper: month one (illustrative)
15. Watch me do it: tests + TS client
16. Recap + try this now

## Lecture transcript

### Capstone part 2

Your content service works on your machine. Now let's make it something twenty five clients can rely on. In this final lesson you'll test it without calling real providers, evaluate every route across languages, add observability, deploy it in a container, and ship a TypeScript client for the agency's web apps. Then you'll assemble the launch pack.

### Why fakes?

Why test with fakes? Because you want to check your own logic, like routing, validation, redaction and budgets, without spending money or depending on provider uptime. Think of a flight simulator. Pilots practice engine failures in a simulator because you can't wait for a real one. A fake provider is your simulator: it can succeed, fail, or return broken JSON on demand.

### Four tests

The lesson includes four tests. The happy path: a fake provider returns good JSON and the service returns content. Fallback: the primary fake raises a retryable error, and the response shows the fallback provider. Invalid JSON: the fake returns plain text and the service returns a five oh two instead of broken content. And redaction: a spy provider records what it received, and the test confirms the email address in the brief never reached it.

### Eval harness

Step two is an evaluation harness. Write thirty to sixty briefs per task across English, Arabic and Urdu. Score outputs on brand voice, fidelity to the brief with no invented claims, length limits, language correctness, and whether needs review is set when it should be. For example, a brief claiming clinically proven without evidence should be flagged. Combine a calibrated judge with code checks, like counting hashtags and requiring the ad tag when content is sponsored. Run every route's primary and fallback.

### Simple eval example

A simple example from the eval. A brief for a skincare cream says: our customers love it, it reduces wrinkles. The rubric checks that the output doesn't invent a percentage or a clinical claim, and that needs review is true, with a reason like, performance claim needs evidence. If a model writes, clinically proven to reduce wrinkles by forty percent, the case fails, no matter how beautiful the prose is.

### Observability

Step three is observability. Write structured logs with the request id, tenant, task, provider, model, tokens, cost, latency and outcome, but never the raw brief. Track metrics: requests, errors by class, fallback activations, breaker state, latency percentiles and cost per tenant per day. Add traces around provider calls. And set alerts on error rate, fallback rate, spend anomalies and validation failures.

### Reading eval results

Here's how to read your eval results. Look at each task, language and provider combination separately, because averages hide problems. If Arabic social posts score lower on the fallback than the primary, that fallback may need an Arabic specific prompt, or shouldn't be used for Arabic at all. If a model rarely sets needs review on risky briefs, strengthen the instruction or add a code check. And look at cost per item next to quality, because the best route is the cheapest one that clears your quality bar.

### Deployment

Step four is deployment. The Dockerfile installs requirements, runs as a non root user and starts uvicorn with a couple of workers. Deploy behind TLS on a container platform in the region your data agreements require. Inject provider keys or workload identity from a secret manager. Run at least two replicas with shared budget state. And set timeouts above your longest expected generation, or use the streaming endpoint for long drafts.

### TypeScript client

Step five is a TypeScript client for the agency's web apps. The generate function posts to your service with the caller's key, maps a four twenty nine to a friendly budget message and a five oh three to try again shortly, and returns typed content. A stream draft helper consumes your server sent events for long drafts. Remember: the browser never sees provider keys, and for server rendered apps, call the service from your server. For Arabic and Urdu, set the output container's direction to right to left.

### Business story: the first month

Now the realistic business story, with illustrative numbers. In week one, two clients run in shadow mode, and the eval flags awkward Urdu hashtags, fixed with a language specific rule. In week two, five clients go live, and a provider incident opens the circuit breaker for twelve minutes; the fallback handles traffic and clients notice nothing. In week three, the cost report shows social posts are cheaper on the fallback at equal quality, so the team swaps the route after rerunning the eval. In week four, all twenty five clients are live, with per client AI costs invoiced monthly.

### Runbook essentials

Step six is the runbook. Document how to switch a task's primary model: edit the routes file, run the eval, deploy. How to force a fallback during an incident. How to raise a tenant's budget. How to rotate keys. And who's on call. Pair it with the production checklist from the security lesson, with evidence for each row. That combination turns a clever project into a dependable service.

### Common mistakes

Common mistakes at this stage. Tests that call real providers, which makes them slow, flaky and expensive. Evals run only in English. Logs that include raw briefs with personal data. And a browser client holding provider keys. Each one undermines the careful work you've done.

### Adding a new task

One more question you'll face after launch: how do you add a fourth task, like blog outlines? Add the task to the schemas, write its rules, add a route in the configuration with a primary and fallback, write thirty eval briefs across your languages, and run the eval before exposing it in the client. Because the architecture separates configuration, schemas and evaluation, new tasks become a routine change rather than a new project.

### Deeper: month one (illustrative)

Let's deepen the first month's story with illustrative numbers. In shadow mode, the eval showed Urdu hashtags were often awkward transliterations, so the team added a rule: Urdu posts use English hashtags only. During the provider incident in week two, the breaker opened for twelve minutes and about two hundred requests went to the fallback; the eval had already shown the fallback's English quality was equal and its Arabic slightly lower, so they temporarily routed Arabic requests to a queue for the few minutes it mattered. In week three, the cost report showed the fallback was cheaper for social posts at equal quality, so the routes swapped after a fresh eval run. By week four, the agency invoiced AI usage per client for the first time.

### Watch me do it: tests + TS client

Watch me do it. Let's run the test suite and the client. The fake provider returns good JSON, or raises a retryable error, or returns plain text, depending on how I construct it. The test client posts the same body each time, which includes an email address in the brief. Happy path: I replace Anthropic with a fake and expect status two hundred and the title blue pottery mug. Fallback: Anthropic fails, Google succeeds, and the response says the provider was Google. Invalid JSON: the fake returns not json and the service returns five oh two. Redaction: a spy provider records the brief it received, and the test asserts the email isn't in it. I run pytest: four passes, no API calls. Then the TypeScript client: generate posts to slash v1 slash generate, maps four twenty nine and five oh three to friendly messages, and returns typed content. I call it from a small page, and the title and body render, with right to left direction for Arabic.

### Recap + try this now

Quick recap, and congratulations. You've tested with fakes, evaluated routes across languages, added observability, deployed securely and shipped a TypeScript client and a runbook. Across the course you've learned the provider landscape, keys, message formats, streaming, tools, structured outputs, multimodal inputs, embeddings, batch, caching, reliability, cost tracking, abstraction, routing, cloud platforms, open models and security. Try this now: add the four tests, run the eval across two providers and three languages, deploy to staging, build a tiny web page with the client, write the runbook, and then take the final exam.

## Key takeaways

- Test service logic with fake providers: happy path, fallback, invalid JSON and redaction.
- Evaluate every route's primary and fallback across tasks and languages with rubrics and code checks.
- Observe requests, errors, fallbacks, costs and latency with structured logs, metrics, traces and alerts.
- Deploy in containers behind TLS with secrets from a manager and shared budget state.
- Ship a typed TypeScript client that never exposes provider keys, plus a runbook.

## Try it

Add the four tests, run the eval across two providers and three languages, deploy to a staging environment, build a tiny web page using the TypeScript client, and write the runbook.

- [Previous: Capstone part 1: build a provider-agnostic content-generation microservice](https://optimizeall.com/learn/ai-platform-apis-integration/capstone-build-content-microservice)
- [All lessons of Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API](https://optimizeall.com/learn/ai-platform-apis-integration)
