Building Production AI AgentsProduction engineering: cost, deployment and frameworks · Lesson 16 of 18

The agent SDK and framework landscape in 2026

Article · 15 min · 9 min lecture

Video lecture

The agent SDK and framework landscape in 2026

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

The agent framework landscape

  • Do you need one?
  • The major options
  • How to choose

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Do you need a framework?

The loop from lesson 3 plus the patterns in this course can take you surprisingly far. Frameworks earn their place when they save you from rebuilding durable state, approvals, tracing, multi-agent coordination or sandboxes. Choose based on what you would otherwise have to build, your language, your model providers, and your deployment target. And keep your business logic (tools, prompts, evals) portable.

The landscape (as of September 2026; verify current docs)

Provider SDKs and harnesses

  • Anthropic API tool runner: part of the regular Anthropic SDKs (beta); runs the tool loop for tools you define, with per-turn hooks. Minimal and flexible.
  • Claude Agent SDK (claude-agent-sdk for Python, @anthropic-ai/claude-agent-sdk for TypeScript): Claude Code packaged as a library. Built-in tools for files, shell and web, sub-agents, hooks, permission modes, sessions and MCP support. Strong for coding, file and ops agents you host yourself.
  • Claude Managed Agents (beta): Anthropic runs the loop and a per-session container; you define versioned agent configs and start sessions.
  • OpenAI Agents SDK (openai-agents, plus a JS/TS version): Agents, Runner, tools (functions, MCP, hosted tools), handoffs, agents-as-tools, guardrails, human-in-the-loop approvals, sessions, tracing, plus sandbox, realtime and voice agents. Described by OpenAI as provider-agnostic.
  • Google Agent Development Kit (ADK) (google-adk; also Java, Go, TypeScript and Kotlin): code-first agents, a graph-based workflow runtime, multi-agent hierarchies, MCP and OpenAPI tools, tool confirmation for HITL, evaluation, and deployment to Cloud Run or Google's managed agent runtime. Optimized for Gemini but model-agnostic.

Independent frameworks

  • LangGraph (LangChain): low-level orchestration for long-running, stateful agents as graphs; durable execution with checkpointers, interrupts for HITL, memory; pairs with LangSmith for tracing/evals and deployment. "Deep Agents" is a higher-level package on top.
  • CrewAI: role-based "Crews" of agents plus event-driven "Flows" for precise control; popular for business automation prototypes.
  • Microsoft Agent Framework: the successor that unifies AutoGen and Semantic Kernel (1.0 released in April 2026 for Python and .NET) with MCP and A2A support. AutoGen itself is in maintenance mode, so new projects should not start on it.

Other notable options include Pydantic AI, LlamaIndex agents, Mastra (TypeScript), and cloud agent services from AWS (Bedrock AgentCore) and Microsoft Foundry.

A comparison lens

NeedStrong fits
Coding/file/ops agent with built-in toolsClaude Agent SDK
Customer-facing multi-agent with handoffs and guardrailsOpenAI Agents SDK
Complex stateful graphs, durable HITLLangGraph
Gemini/Google Cloud-centric, multi-language teamsGoogle ADK
.NET or Microsoft-centric enterpriseMicrosoft Agent Framework
Fast role-based business automation prototypesCrewAI
Hosted loop + sandbox, minimal infraManaged agent platforms
Maximum control, minimal dependenciesYour own loop + provider SDK

Evaluation criteria for your shortlist

  1. Transparency: can you see and version every prompt and tool call?
  2. Model flexibility: can you swap providers without rewrites?
  3. State and HITL: checkpoints, interrupts, resumable approvals?
  4. Observability: OpenTelemetry export or equivalent?
  5. Security: permission model, sandboxing, secrets handling.
  6. MCP support for tool integrations.
  7. Maturity: release cadence, breaking changes, maintenance status, license.
  8. Deployment fit: your cloud, regions and data residency.

Worked example: choosing for a PK/UK SaaS company

A Lahore- and London-based SaaS company wants (a) an internal ops agent that reads logs and files tickets, and (b) a customer onboarding assistant with handoffs to billing and support.

  • Ops agent → Claude Agent SDK (built-in file/shell tools, hooks to block destructive commands, MCP for the ticketing system), hosted in their own containers.
  • Onboarding assistant → OpenAI Agents SDK (handoffs, guardrails, sessions) or LangGraph if they need complex durable approvals; they prototype both against the same 30-case eval.
  • Shared layer: tools exposed as MCP servers so either framework can use them; OpenTelemetry to one backend; prompts in their own repo.

The portable layer (MCP tools, evals, traces) means a future framework switch costs days, not months.

Hands-on: the same tiny agent in two SDKs

Claude Agent SDK (Python):

import anyio
from claude_agent_sdk import query, ClaudeAgentOptions, AssistantMessage, TextBlock

async def main():
    options = ClaudeAgentOptions(
        system_prompt="You are an ops assistant. Summarize errors; never modify files.",
        allowed_tools=["Read", "Grep", "Glob"],     # auto-approved; others need permission
        disallowed_tools=["Bash", "Write", "Edit"],
        max_turns=8,
        cwd="./logs",
    )
    async for message in query(prompt="Find the top 3 error types in today's logs.", options=options):
        if isinstance(message, AssistantMessage):
            for block in message.content:
                if isinstance(block, TextBlock):
                    print(block.text)

anyio.run(main)

OpenAI Agents SDK (Python):

from agents import Agent, Runner, function_tool

@function_tool
def count_errors(date: str) -> dict:
    """Return counts of error types for a date (YYYY-MM-DD)."""
    return {"timeout": 42, "auth_failed": 17, "db_lock": 5}   # replace with a real query

agent = Agent(name="Ops assistant",
              instructions="Summarize errors. Never take actions.",
              tools=[count_errors])

result = Runner.run_sync(agent, "What were the top 3 error types on 2026-09-24?")
print(result.final_output)

Both require API keys in environment variables (ANTHROPIC_API_KEY, OPENAI_API_KEY). Check each SDK's current docs for model selection and options.

Pitfalls

  • Choosing by GitHub stars rather than your eval results.
  • Starting new projects on frameworks in maintenance mode.
  • Framework lock-in of tools and prompts; keep them portable.
  • Upgrading SDK major versions without re-running evals.

Measuring success

For a shortlist of two, build the same agent against the same eval and compare pass^k, cost per task, latency, lines of your own code, and how easily you can see and change prompts.

Key takeaways

  • Frameworks should replace infrastructure you'd otherwise build; keep tools, prompts and evals portable.
  • Provider SDKs: Anthropic tool runner, Claude Agent SDK, Managed Agents, OpenAI Agents SDK, Google ADK.
  • LangGraph, CrewAI and Microsoft Agent Framework (successor to AutoGen and Semantic Kernel) are major independent options.
  • Evaluate on transparency, model flexibility, state/HITL, observability, security, MCP support and maturity.
  • Decide with a bake-off on your own eval set, not popularity.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A team plans a new multi-agent project on AutoGen in 2026. What should you advise?
  2. Which option packages Claude Code's harness with built-in file, shell and web tools as a library you host?
  3. What makes a later framework switch cheap?

Put it into practice

Shortlist two frameworks for your agent using the eight criteria. Build the smallest version in each and compare them on five eval cases.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.