Building Production AI AgentsProduction engineering: cost, deployment and frameworks · Lesson 15 of 18

Deploying agents: architectures, sandboxes and safe rollout

Article · 16 min · 9 min lecture

Video lecture

Deploying agents: architectures, sandboxes and safe rollout

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Deploying agents

  • Reference architecture
  • Hosting options
  • Sandboxing
  • Safe rollout

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

From script to service

A production agent is a small distributed system: an API that accepts tasks, a worker that runs the loop, a state store, a queue, tool integrations with credentials, sandboxes for code or browsing, an approval service, tracing, and configuration (prompts, models, limits) that changes independently of code.

Reference architecture

Client (web app / Slack / API)
   │  POST /runs  → returns run_id
   ▼
API service ──► Runs DB (state, checkpoints, approvals, audit)
   │                 ▲
   ▼                 │
Queue ──► Agent workers (loop, guardrails, budgets) ──► Model APIs
                     │         │
                     │         ├──► Tool gateway (per-user credentials, allowlists, rate limits)
                     │         └──► Sandbox (code/browser, no secrets, egress allowlist)
                     ▼
               Tracing / metrics / alerts

Key design choices:

  • Stateless workers, stateful store: any worker can resume any run (lesson 9).
  • Tool gateway: one place to hold credentials, enforce allowlists, rate limits and audit. Increasingly this is where MCP servers sit (see the MCP course).
  • Config as data: prompts, model IDs, effort levels and budgets stored with versions, so you can roll forward and back without redeploying.

Hosting options

OptionFitsWatch out for
Serverless functionsShort, bursty, stateless stepsExecution time limits; cold starts
Containers (Kubernetes, Cloud Run, ECS, Azure Container Apps)Long-running workers, custom depsScaling queues, cost of idle
Workflow / durable execution enginesLong, multi-step runs with retriesLearning curve
Managed agent platformsYou want the provider to run the loop and sandboxFeature and region availability, lock-in, data residency
Framework deployment platforms (e.g., LangSmith Deployment, Google's agent runtimes)Framework-native agentsCoupling to one framework

Anthropic's Claude Managed Agents (beta at the time of writing) run the loop and a per-session container on Anthropic's infrastructure; the Claude Agent SDK gives you Claude Code's harness to host yourself. OpenAI's Agents SDK offers sandbox agents with local or hosted sandbox clients. Google's ADK deploys to Cloud Run or Google's managed agent engine. Check each vendor's current availability and data-processing terms.

Sandboxing agents that run code or browse

  • Isolated container or micro-VM per session; destroy after.
  • No production secrets inside; tools that need credentials run outside the sandbox via the gateway.
  • Egress allowlist; CPU, memory, disk and time limits.
  • Read-only mounts by default; explicit output directory.
  • Scan artifacts leaving the sandbox (files, links) before users download them.

Versioning and safe rollout

Treat each change to prompts, tools, models or limits as a release:

  1. Eval gate in CI (lesson 12): no merge if pass^k drops or cost rises beyond threshold.
  2. Shadow mode: run the new version on real inputs without acting; compare outputs.
  3. Canary: route 5% of traffic, watch dashboards, then ramp.
  4. Feature flags per tenant for risky capabilities.
  5. Instant rollback by switching config version.
  6. Model deprecation plan: providers retire model versions; keep model IDs in config, run your eval suite on successors early, and track provider deprecation notices.

Worked example: rolling out a support agent at a Saudi telecom reseller

  • Phase 1 (shadow, 2 weeks): agent drafts replies for every ticket; agents compare with human replies; eval set grows from disagreements.
  • Phase 2 (assist): drafts shown to human agents, who send or edit; edit rate tracked.
  • Phase 3 (canary autonomy): for two low-risk intents (balance queries, SIM activation status), the agent replies directly to 5% of customers, with instant rollback via flag.
  • Phase 4 (ramp): expanded to 50% for those intents after stable metrics; refunds remain human-only.
  • Data residency: hosting and provider regions chosen to meet the company's KSA data requirements (checked with legal against PDPL and sector rules).

Hands-on: a minimal run API with FastAPI and a background worker

# pip install fastapi uvicorn
import uuid, threading
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel

app = FastAPI()
RUNS: dict[str, dict] = {}          # replace with a database + real queue in production

class RunRequest(BaseModel):
    goal: str
    tenant_id: str

def worker(run_id: str):
    run = RUNS[run_id]
    try:
        run["status"] = "running"
        run["result"] = run_agent(run["goal"])      # your loop from lesson 3, with checkpoints
        run["status"] = "succeeded"
    except Exception as exc:
        run["status"], run["error"] = "failed", str(exc)[:500]

@app.post("/runs", status_code=202)
def create_run(req: RunRequest):
    run_id = str(uuid.uuid4())
    RUNS[run_id] = {"goal": req.goal, "tenant": req.tenant_id, "status": "queued"}
    threading.Thread(target=worker, args=(run_id,), daemon=True).start()   # use a queue in prod
    return {"run_id": run_id, "status_url": f"/runs/{run_id}"}

@app.get("/runs/{run_id}")
def get_run(run_id: str):
    run = RUNS.get(run_id)
    if not run:
        raise HTTPException(404, "run not found")
    return {k: v for k, v in run.items() if k != "goal"}

This shows the contract (202 + status URL); swap the thread for a proper queue worker and the dict for your runs table before production.

Pitfalls

  • Long synchronous endpoints that time out behind load balancers.
  • Hard-coded model IDs scattered through code.
  • Sandboxes with production credentials mounted "temporarily".
  • Rolling out autonomy to all traffic at once.

Measuring success

Deployment frequency with no incident, change failure rate, time to rollback, canary metric deltas, and cost/quality per version.

Key takeaways

  • A production agent is a distributed system: API, queue, workers, state, tool gateway, sandbox, tracing.
  • Keep workers stateless, state durable, and prompts/models/limits as versioned config.
  • Sandboxes need isolation, no secrets, egress allowlists and resource limits.
  • Roll out with eval gates, shadow mode, canaries, flags and instant rollback.
  • Plan for model deprecations by keeping model IDs in config and evaluating successors early.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why store prompts, model IDs and budgets as versioned configuration?
  2. Which rollout step runs a new agent version on real inputs without taking actions?
  3. A code-execution sandbox needs to call the CRM. Where should the CRM credential live?

Put it into practice

Draw your agent's deployment architecture using the reference diagram, and write a four-phase rollout plan with the metrics that gate each phase.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.