Building Production AI AgentsProduction engineering: cost, deployment and frameworks · Lesson 15 of 18
Deploying agents: architectures, sandboxes and safe rollout
Video lecture
Deploying agents: architectures, sandboxes and safe rollout
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Deploying agents
Getting an agent to work on your laptop is the easy part. Running it reliably for thousands of users, safely, with the ability to roll back in seconds, is where engineering really starts. In this lesson you'll learn a reference architecture, your hosting options, how to sandbox agents that run code, and how to roll out changes safely.
0:25 Why it matters
Why does deployment architecture matter so much for agents? Because agents break the assumptions of normal web apps. Requests can run for minutes, they call external systems with real side effects, and they sometimes execute code. Think of the difference between a food truck and a restaurant kitchen. A food truck works for a few quick orders. A restaurant needs stations, tickets, a pass where dishes are checked, and a way to handle a rush. Production agents need the restaurant kitchen.
1:00 Reference architecture
Here's the reference architecture. A client, like a web app or Slack, posts a task and gets back a run id. The API stores the run and puts it on a queue. Stateless workers pick it up and run the loop with guardrails and budgets, calling model APIs. Tools go through a gateway that holds credentials, enforces allowlists and rate limits, and audits everything. Code and browsing run in a sandbox with no secrets. And everything emits traces and metrics.
1:35 Three design choices
Three design choices make this work. Stateless workers with a stateful store, so any worker can resume any run. A tool gateway, one place for credentials, allowlists, rate limits and audit, which is increasingly where MCP servers live. And configuration as data: prompts, model ids, effort levels and budgets stored with versions, so you can roll forward and back without redeploying code.
2:02 Hosting options
Hosting options. Serverless functions suit short, bursty steps, but watch time limits and cold starts. Containers suit long running workers. Durable execution engines handle long, multi step runs with retries. Managed agent platforms run the loop and sandbox for you. For example, Anthropic's managed agents, in beta at the time of writing, host the loop and a container per session, while the Claude Agent SDK gives you Claude Code's harness to host yourself. Framework platforms deploy framework native agents. Check each vendor's current regions and data terms.
2:40 Simple example: laptop to service
A simple example. A small agency runs a content agent from a single script on a laptop. It works until a teammate closes the laptop mid run, and two clients get half finished drafts. The next version takes requests through a tiny API that returns a run id, puts work on a queue, and processes it in a container that saves progress after every step. The laptop can close, the container can restart, and nobody loses work. It's not fancy, but it's the first real deployment.
3:17 Sandbox rules
If your agent runs code or browses the web, sandbox it properly. Use an isolated container or micro VM per session and destroy it afterwards. Keep production secrets out; anything needing credentials runs outside through the gateway. Add an egress allowlist and limits on CPU, memory, disk and time. Mount read only by default. And scan any files or links leaving the sandbox before users download them.
3:46 Safe rollout
Now safe rollout. Treat every change to prompts, tools, models or limits as a release. Gate merges on your eval suite. Run new versions in shadow mode on real inputs without acting. Canary to a small slice of traffic, watch the dashboards, then ramp. Use feature flags per tenant for risky capabilities, and make rollback a config switch. And plan for model deprecations: keep model ids in config and run your evals on successor models early.
4:19 Example: KSA telecom support agent
A worked example. A telecom reseller in Saudi Arabia rolled out a support agent in four phases. First, two weeks of shadow mode, drafting replies that staff compared with their own. Second, assist mode, where staff send or edit drafts while the edit rate is tracked. Third, a canary: for two low risk intents, balance queries and SIM activation status, the agent replies directly to five percent of customers. Fourth, a ramp to half of those intents once metrics held steady. Refunds stayed human only. Hosting and provider regions were chosen with legal to meet Saudi data requirements.
5:02 Hands-on: run API contract
The hands on code sketches the run API contract with FastAPI. Posting a goal returns a two oh two with a run id and a status URL, and a worker processes it in the background. Getting the run returns its status and result. The sketch uses a thread and a dictionary to stay short. In production, swap those for a real queue and your runs table. Avoid the classic pitfalls: long synchronous endpoints, model ids hard coded everywhere, sandboxes with production credentials, and turning on autonomy for all traffic at once.
5:42 Rolling back
How do you roll back an agent? Because prompts, model ids and limits live in versioned configuration, rollback is usually switching the active config version, not redeploying code. Practice it before you need it: in staging, promote a new config, run a few requests, then roll back and confirm behavior returns to the previous version. Also decide in advance what triggers a rollback, like success rate dropping below a threshold or cost per run jumping, so nobody debates it during an incident.
6:18 Deeper: the KSA telecom rollout (illustrative)
Let's deepen the Saudi telecom reseller rollout. In shadow mode, the agent's drafts matched what human agents actually sent in a majority of balance and SIM questions, illustrative numbers, but disagreed often on anything involving refunds, which confirmed refunds should stay human. In assist mode, edit rates fell week by week as prompts improved. The canary at five percent ran for ten days with customer satisfaction scores equal to the human baseline. One incident happened: a billing API outage made the agent reply with, I can't find your balance. The feature flag switched those intents back to human agents in under a minute, which convinced the operations director that autonomy could be trusted because it could be reversed.
7:09 Watch me do it: the run API
Watch me do it. Let's walk through the run API. The request model has a goal and a tenant id. Create run makes a new run id, stores the goal, tenant and a queued status, starts a background worker, and returns status two oh two with the run id and a status URL, so the client never waits on a long request. The worker marks the run as running, calls run agent, and stores the result with a succeeded status, or catches any exception and stores a failed status with a short error message. Get run looks the id up, returns four oh four if it's missing, and returns everything except the goal. I'll post a goal with curl, get back a run id, and poll the status URL: queued, running, succeeded. In production, the thread becomes a queue worker and the dictionary becomes the runs table from the durability lesson.
8:15 Try this now
Try this now. Draw your agent's deployment on one page using the reference architecture: where requests come in, where state lives, where tools and credentials sit, and where code runs. Circle the single point that would hurt most if it failed. Then write a four phase rollout plan: shadow, assist, canary and ramp, with one metric that must hold before each phase moves forward. Share it with whoever would be paged at night.
8:47 Recap
To recap: design a small distributed system with stateless workers, durable state, a tool gateway and proper sandboxes. Keep configuration versioned, and roll out with eval gates, shadow mode, canaries, flags and instant rollback. Your next step: draw your own deployment architecture from the reference diagram and write a four phase rollout plan with the metrics that gate each phase.
From script to service
A production agent is a small distributed system: an API that accepts tasks, a worker that runs the loop, a state store, a queue, tool integrations with credentials, sandboxes for code or browsing, an approval service, tracing, and configuration (prompts, models, limits) that changes independently of code.
Reference architecture
Client (web app / Slack / API)
│ POST /runs → returns run_id
▼
API service ──► Runs DB (state, checkpoints, approvals, audit)
│ ▲
▼ │
Queue ──► Agent workers (loop, guardrails, budgets) ──► Model APIs
│ │
│ ├──► Tool gateway (per-user credentials, allowlists, rate limits)
│ └──► Sandbox (code/browser, no secrets, egress allowlist)
▼
Tracing / metrics / alertsKey design choices:
- Stateless workers, stateful store: any worker can resume any run (lesson 9).
- Tool gateway: one place to hold credentials, enforce allowlists, rate limits and audit. Increasingly this is where MCP servers sit (see the MCP course).
- Config as data: prompts, model IDs, effort levels and budgets stored with versions, so you can roll forward and back without redeploying.
Hosting options
| Option | Fits | Watch out for |
|---|---|---|
| Serverless functions | Short, bursty, stateless steps | Execution time limits; cold starts |
| Containers (Kubernetes, Cloud Run, ECS, Azure Container Apps) | Long-running workers, custom deps | Scaling queues, cost of idle |
| Workflow / durable execution engines | Long, multi-step runs with retries | Learning curve |
| Managed agent platforms | You want the provider to run the loop and sandbox | Feature and region availability, lock-in, data residency |
| Framework deployment platforms (e.g., LangSmith Deployment, Google's agent runtimes) | Framework-native agents | Coupling to one framework |
Anthropic's Claude Managed Agents (beta at the time of writing) run the loop and a per-session container on Anthropic's infrastructure; the Claude Agent SDK gives you Claude Code's harness to host yourself. OpenAI's Agents SDK offers sandbox agents with local or hosted sandbox clients. Google's ADK deploys to Cloud Run or Google's managed agent engine. Check each vendor's current availability and data-processing terms.
Sandboxing agents that run code or browse
- Isolated container or micro-VM per session; destroy after.
- No production secrets inside; tools that need credentials run outside the sandbox via the gateway.
- Egress allowlist; CPU, memory, disk and time limits.
- Read-only mounts by default; explicit output directory.
- Scan artifacts leaving the sandbox (files, links) before users download them.
Versioning and safe rollout
Treat each change to prompts, tools, models or limits as a release:
- Eval gate in CI (lesson 12): no merge if pass^k drops or cost rises beyond threshold.
- Shadow mode: run the new version on real inputs without acting; compare outputs.
- Canary: route 5% of traffic, watch dashboards, then ramp.
- Feature flags per tenant for risky capabilities.
- Instant rollback by switching config version.
- Model deprecation plan: providers retire model versions; keep model IDs in config, run your eval suite on successors early, and track provider deprecation notices.
Worked example: rolling out a support agent at a Saudi telecom reseller
- Phase 1 (shadow, 2 weeks): agent drafts replies for every ticket; agents compare with human replies; eval set grows from disagreements.
- Phase 2 (assist): drafts shown to human agents, who send or edit; edit rate tracked.
- Phase 3 (canary autonomy): for two low-risk intents (balance queries, SIM activation status), the agent replies directly to 5% of customers, with instant rollback via flag.
- Phase 4 (ramp): expanded to 50% for those intents after stable metrics; refunds remain human-only.
- Data residency: hosting and provider regions chosen to meet the company's KSA data requirements (checked with legal against PDPL and sector rules).
Hands-on: a minimal run API with FastAPI and a background worker
# pip install fastapi uvicorn
import uuid, threading
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
app = FastAPI()
RUNS: dict[str, dict] = {} # replace with a database + real queue in production
class RunRequest(BaseModel):
goal: str
tenant_id: str
def worker(run_id: str):
run = RUNS[run_id]
try:
run["status"] = "running"
run["result"] = run_agent(run["goal"]) # your loop from lesson 3, with checkpoints
run["status"] = "succeeded"
except Exception as exc:
run["status"], run["error"] = "failed", str(exc)[:500]
@app.post("/runs", status_code=202)
def create_run(req: RunRequest):
run_id = str(uuid.uuid4())
RUNS[run_id] = {"goal": req.goal, "tenant": req.tenant_id, "status": "queued"}
threading.Thread(target=worker, args=(run_id,), daemon=True).start() # use a queue in prod
return {"run_id": run_id, "status_url": f"/runs/{run_id}"}
@app.get("/runs/{run_id}")
def get_run(run_id: str):
run = RUNS.get(run_id)
if not run:
raise HTTPException(404, "run not found")
return {k: v for k, v in run.items() if k != "goal"}This shows the contract (202 + status URL); swap the thread for a proper queue worker and the dict for your runs table before production.
Pitfalls
- Long synchronous endpoints that time out behind load balancers.
- Hard-coded model IDs scattered through code.
- Sandboxes with production credentials mounted "temporarily".
- Rolling out autonomy to all traffic at once.
Measuring success
Deployment frequency with no incident, change failure rate, time to rollback, canary metric deltas, and cost/quality per version.
Key takeaways
- A production agent is a distributed system: API, queue, workers, state, tool gateway, sandbox, tracing.
- Keep workers stateless, state durable, and prompts/models/limits as versioned config.
- Sandboxes need isolation, no secrets, egress allowlists and resource limits.
- Roll out with eval gates, shadow mode, canaries, flags and instant rollback.
- Plan for model deprecations by keeping model IDs in config and evaluating successors early.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Draw your agent's deployment architecture using the reference diagram, and write a four-phase rollout plan with the metrics that gate each phase.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.