Building Production AI AgentsProduction engineering: cost, deployment and frameworks · Lesson 16 of 18
The agent SDK and framework landscape in 2026
Video lecture
The agent SDK and framework landscape in 2026
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 The agent framework landscape
There are more agent frameworks than anyone can keep up with, and new ones every month. The question isn't which one is best. It's which one saves you the most work for your specific agent, without locking you in. In this lesson you'll get a clear map of the major SDKs and frameworks as of September twenty twenty six, and a method for choosing.
0:28 Why it matters
Why does framework choice matter? Because it's easy to reverse on day one and expensive to reverse on day one hundred. Frameworks shape how you write tools, store state and debug. Think of choosing a phone ecosystem. Switching later means moving photos, apps and habits. The smart move is to keep your most valuable things portable, like your photos in a format any phone reads. For agents, the portable things are your tools, prompts, evals and traces. Keep those independent and you can switch frameworks without starting over.
1:06 Do you need a framework?
First, do you need a framework at all? Your own loop plus the patterns in this course go a long way. Frameworks earn their place when they save you from rebuilding durable state, approvals, tracing, multi agent coordination or sandboxes. Choose based on what you'd otherwise build, your language, your model providers and where you deploy. Whatever you pick, keep your business logic, meaning tools, prompts and evals, portable.
1:36 Provider SDKs
Start with the provider SDKs. Anthropic's API has a tool runner that handles the loop for tools you define. The Claude Agent SDK is Claude Code packaged as a library, with built in file, shell and web tools, sub agents, hooks and permissions, which you host yourself. Claude Managed Agents, in beta, runs the loop and a container for you. OpenAI's Agents SDK offers agents, a runner, handoffs, agents as tools, guardrails, approvals, sessions, tracing, and sandbox and voice agents. Google's Agent Development Kit brings code first agents, a graph workflow runtime, tool confirmation and deployment to Google Cloud.
2:19 Independent frameworks
Then the independent frameworks. LangGraph, from LangChain, models long running, stateful agents as graphs, with checkpointers for durable execution and interrupts for human review, and pairs with LangSmith for tracing and evals. CrewAI offers role based crews of agents plus event driven flows. And Microsoft Agent Framework unifies AutoGen and Semantic Kernel, with version one released in April twenty twenty six for Python and dot net. AutoGen itself is now in maintenance mode, so don't start new projects on it.
2:54 Simple example: two developers
A simple example. A solo developer wants an agent that tidies their downloads folder and summarizes what changed. They need file tools, safe permissions and not much else. The Claude Agent SDK gives them file and search tools out of the box, and they can disallow deletion and shell commands in a few lines. A different developer wants a customer chat with a triage agent handing off to billing and support. The OpenAI Agents SDK's handoffs and guardrails fit that shape naturally. Different jobs, different best fits.
3:32 Eight selection criteria
To choose, score your shortlist on eight criteria. Transparency: can you see every prompt and tool call? Model flexibility: can you swap providers? State and human review: checkpoints, interrupts, resumable approvals? Observability: can you export traces? Security: permissions, sandboxing and secrets. MCP support for tools. Maturity: release cadence, breaking changes, maintenance status and license. And deployment fit: your cloud, regions and data residency.
3:59 Example: two agents, two choices
A worked example. A software company with teams in Lahore and London needs two agents. An internal ops agent that reads logs and files tickets, and a customer onboarding assistant with handoffs to billing and support. For ops, they pick the Claude Agent SDK for its built in file tools, with hooks blocking destructive commands. For onboarding, they prototype the OpenAI Agents SDK and LangGraph against the same thirty case eval. Crucially, their tools are MCP servers any framework can use, traces go to one OpenTelemetry backend, and prompts live in their own repo. A future switch costs days, not months.
4:43 Hands-on: same agent, two SDKs
The hands on section builds the same tiny ops agent in two SDKs. With the Claude Agent SDK, you call query with options that auto approve read, grep and glob, disallow bash and file writes, cap the turns and set the working folder. With the OpenAI Agents SDK, you decorate a Python function as a tool, create an agent with instructions, and run it with the runner. Same idea, different ergonomics. Compare how easily you can see and change what's happening.
5:18 Traps and the bake-off
Avoid the common traps: choosing by popularity instead of your eval results, starting new work on frameworks in maintenance mode, locking tools and prompts into one framework's format, and upgrading major versions without re running evals. The decisive test is a bake off: build the same agent in two options, run the same eval, and compare reliability, cost, latency, how much code you wrote, and how easily you can see and change prompts.
5:50 If a framework stalls
What if you pick a framework and it stalls or changes direction? That's why the portable layer matters. If your tools are MCP servers, your prompts live in your repository, your evals run independently, and your traces use open standards, then moving frameworks means rewriting the orchestration glue, not the valuable parts. Check a framework's release history, maintainers and issue activity before adopting, and revisit your choice once a year with the same bake off method.
6:23 Deeper: the SaaS bake-off (illustrative)
Let's deepen the Lahore and London SaaS example. For the ops agent, the Claude Agent SDK gave them file and search tools on day one, and a pre tool use hook blocked any shell command containing delete or drop. For onboarding, the bake off on thirty cases showed both candidates within a few points on quality, illustrative figures, but LangGraph's checkpointing made their multi day approval flow simpler, while the OpenAI Agents SDK's handoffs made the triage conversation cleaner. They chose the Agents SDK for the chat and kept one LangGraph graph for the approval workflow. It sounds messy, but because tools were MCP servers and traces went to one backend, both parts shared everything that mattered.
7:14 Watch me do it: two SDKs
Watch me do it. Let's compare the two snippets line by line. In the Claude Agent SDK version, I import query and the options class. The options set a system prompt that forbids modifying files, auto approve read, grep and glob, explicitly disallow bash, write and edit, cap the run at eight turns and set the working folder to the logs directory. Then I loop over the messages query streams back and print the text blocks from assistant messages. In the OpenAI Agents SDK version, I decorate count errors as a function tool, so its docstring and type hints become the schema. I create an agent with a name, instructions and the tool, then call runner run sync and print the final output. Same job, different ergonomics: one gives me built in file tools and permissions, the other makes custom tools and handoffs very quick.
8:17 Try this now
Try this now. Score two frameworks you're considering against the eight criteria in the lesson, on a simple one to three scale. Then build the smallest possible version of your agent in both: one tool, one goal, run five eval cases each. Compare reliability, cost, lines of code you wrote, and how easy it was to see the actual prompt and tool calls. Write down your choice and the reason, so the next person understands why.
8:50 Recap
To recap: frameworks should replace infrastructure you'd otherwise build. Know the main provider SDKs and independent frameworks, score them on eight criteria, keep a portable layer, and decide with a bake off. Your next step: shortlist two frameworks for your agent, build the smallest version in each, and compare them on five eval cases.
Do you need a framework?
The loop from lesson 3 plus the patterns in this course can take you surprisingly far. Frameworks earn their place when they save you from rebuilding durable state, approvals, tracing, multi-agent coordination or sandboxes. Choose based on what you would otherwise have to build, your language, your model providers, and your deployment target. And keep your business logic (tools, prompts, evals) portable.
The landscape (as of September 2026; verify current docs)
Provider SDKs and harnesses
- Anthropic API tool runner: part of the regular Anthropic SDKs (beta); runs the tool loop for tools you define, with per-turn hooks. Minimal and flexible.
- Claude Agent SDK (
claude-agent-sdkfor Python,@anthropic-ai/claude-agent-sdkfor TypeScript): Claude Code packaged as a library. Built-in tools for files, shell and web, sub-agents, hooks, permission modes, sessions and MCP support. Strong for coding, file and ops agents you host yourself. - Claude Managed Agents (beta): Anthropic runs the loop and a per-session container; you define versioned agent configs and start sessions.
- OpenAI Agents SDK (
openai-agents, plus a JS/TS version): Agents, Runner, tools (functions, MCP, hosted tools), handoffs, agents-as-tools, guardrails, human-in-the-loop approvals, sessions, tracing, plus sandbox, realtime and voice agents. Described by OpenAI as provider-agnostic. - Google Agent Development Kit (ADK) (
google-adk; also Java, Go, TypeScript and Kotlin): code-first agents, a graph-based workflow runtime, multi-agent hierarchies, MCP and OpenAPI tools, tool confirmation for HITL, evaluation, and deployment to Cloud Run or Google's managed agent runtime. Optimized for Gemini but model-agnostic.
Independent frameworks
- LangGraph (LangChain): low-level orchestration for long-running, stateful agents as graphs; durable execution with checkpointers, interrupts for HITL, memory; pairs with LangSmith for tracing/evals and deployment. "Deep Agents" is a higher-level package on top.
- CrewAI: role-based "Crews" of agents plus event-driven "Flows" for precise control; popular for business automation prototypes.
- Microsoft Agent Framework: the successor that unifies AutoGen and Semantic Kernel (1.0 released in April 2026 for Python and .NET) with MCP and A2A support. AutoGen itself is in maintenance mode, so new projects should not start on it.
Other notable options include Pydantic AI, LlamaIndex agents, Mastra (TypeScript), and cloud agent services from AWS (Bedrock AgentCore) and Microsoft Foundry.
A comparison lens
| Need | Strong fits |
|---|---|
| Coding/file/ops agent with built-in tools | Claude Agent SDK |
| Customer-facing multi-agent with handoffs and guardrails | OpenAI Agents SDK |
| Complex stateful graphs, durable HITL | LangGraph |
| Gemini/Google Cloud-centric, multi-language teams | Google ADK |
| .NET or Microsoft-centric enterprise | Microsoft Agent Framework |
| Fast role-based business automation prototypes | CrewAI |
| Hosted loop + sandbox, minimal infra | Managed agent platforms |
| Maximum control, minimal dependencies | Your own loop + provider SDK |
Evaluation criteria for your shortlist
- Transparency: can you see and version every prompt and tool call?
- Model flexibility: can you swap providers without rewrites?
- State and HITL: checkpoints, interrupts, resumable approvals?
- Observability: OpenTelemetry export or equivalent?
- Security: permission model, sandboxing, secrets handling.
- MCP support for tool integrations.
- Maturity: release cadence, breaking changes, maintenance status, license.
- Deployment fit: your cloud, regions and data residency.
Worked example: choosing for a PK/UK SaaS company
A Lahore- and London-based SaaS company wants (a) an internal ops agent that reads logs and files tickets, and (b) a customer onboarding assistant with handoffs to billing and support.
- Ops agent → Claude Agent SDK (built-in file/shell tools, hooks to block destructive commands, MCP for the ticketing system), hosted in their own containers.
- Onboarding assistant → OpenAI Agents SDK (handoffs, guardrails, sessions) or LangGraph if they need complex durable approvals; they prototype both against the same 30-case eval.
- Shared layer: tools exposed as MCP servers so either framework can use them; OpenTelemetry to one backend; prompts in their own repo.
The portable layer (MCP tools, evals, traces) means a future framework switch costs days, not months.
Hands-on: the same tiny agent in two SDKs
Claude Agent SDK (Python):
import anyio
from claude_agent_sdk import query, ClaudeAgentOptions, AssistantMessage, TextBlock
async def main():
options = ClaudeAgentOptions(
system_prompt="You are an ops assistant. Summarize errors; never modify files.",
allowed_tools=["Read", "Grep", "Glob"], # auto-approved; others need permission
disallowed_tools=["Bash", "Write", "Edit"],
max_turns=8,
cwd="./logs",
)
async for message in query(prompt="Find the top 3 error types in today's logs.", options=options):
if isinstance(message, AssistantMessage):
for block in message.content:
if isinstance(block, TextBlock):
print(block.text)
anyio.run(main)OpenAI Agents SDK (Python):
from agents import Agent, Runner, function_tool
@function_tool
def count_errors(date: str) -> dict:
"""Return counts of error types for a date (YYYY-MM-DD)."""
return {"timeout": 42, "auth_failed": 17, "db_lock": 5} # replace with a real query
agent = Agent(name="Ops assistant",
instructions="Summarize errors. Never take actions.",
tools=[count_errors])
result = Runner.run_sync(agent, "What were the top 3 error types on 2026-09-24?")
print(result.final_output)Both require API keys in environment variables (ANTHROPIC_API_KEY, OPENAI_API_KEY). Check each SDK's current docs for model selection and options.
Pitfalls
- Choosing by GitHub stars rather than your eval results.
- Starting new projects on frameworks in maintenance mode.
- Framework lock-in of tools and prompts; keep them portable.
- Upgrading SDK major versions without re-running evals.
Measuring success
For a shortlist of two, build the same agent against the same eval and compare pass^k, cost per task, latency, lines of your own code, and how easily you can see and change prompts.
Key takeaways
- Frameworks should replace infrastructure you'd otherwise build; keep tools, prompts and evals portable.
- Provider SDKs: Anthropic tool runner, Claude Agent SDK, Managed Agents, OpenAI Agents SDK, Google ADK.
- LangGraph, CrewAI and Microsoft Agent Framework (successor to AutoGen and Semantic Kernel) are major independent options.
- Evaluate on transparency, model flexibility, state/HITL, observability, security, MCP support and maturity.
- Decide with a bake-off on your own eval set, not popularity.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Shortlist two frameworks for your agent using the eight criteria. Build the smallest version in each and compare them on five eval cases.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.