---
title: "Agent frameworks, SDKs and platforms: how to choose"
description: "You do not have to build every loop by hand You have now written a tool loop yourself, which is the best way to understand agents. In production you will…"
url: https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/agent-frameworks-and-sdks
updated: 2026-10-05
---

Latest AI Techniques: RAG, Tool Use, Agents & MCP · Agent loops, multi-agent patterns and memory · lesson 13 of 20 · 13 min

# Agent frameworks, SDKs and platforms: how to choose

## You do not have to build every loop by hand

You have now written a tool loop yourself, which is the best way to understand agents. In production you will usually lean on a framework, SDK or platform. The market is crowded and moves monthly, so this lesson gives you a durable way to choose, a map of the main categories (with current examples, verified at the time of writing) and a portable hello-world in two SDKs.

## Two questions that separate the options

1. **Who supplies the harness?** The harness is the agent loop plus context management, tool execution, retries, approvals and tracing.
2. **Who supplies the deployment?** Where the loop runs, where tools execute (sandboxes), and who operates it.

| Category | Harness | Deployment | Current examples | Good when |
|---|---|---|---|---|
| Raw API + your loop | You | You | Any model API | You need total control or minimal dependencies |
| SDK loop helpers | SDK | You | Anthropic SDK tool runner; OpenAI Agents SDK | Custom tools, your infrastructure, less boilerplate |
| Full agent harness as a library | SDK (batteries included) | You | Claude Agent SDK (the Claude Code harness as a library); OpenAI Agents SDK sandbox agents | File, shell and coding-style agents on your own infra |
| Graph/workflow frameworks | Framework | You | LangGraph, LlamaIndex workflows, Microsoft's agent frameworks, CrewAI | Explicit state machines, multi-step graphs, human checkpoints, model-agnostic stacks |
| Hosted agent platforms | Provider | Provider | Claude Managed Agents (beta); cloud agent services on AWS, Google Cloud and Azure | Long-running, scheduled or sandboxed agents without running infrastructure |
| No-code / low-code agent builders | Platform | Platform | n8n AI Agent node, Zapier Agents, Microsoft Copilot Studio | Business teams, integrations-heavy automations |

Products change quickly (for example, OpenAI has announced the retirement of its drag-and-drop Agent Builder canvas and points code-based users to the Agents SDK), so check vendor docs before committing.

## Selection criteria that matter

- **Control flow:** do you need an explicit graph (predictable, auditable) or model-directed loops (flexible)? Many teams combine them: a graph with an agent node inside.
- **Model portability:** some frameworks are provider-agnostic; others are tied to one vendor's features. Portability costs you some native features; lock-in costs you negotiating power.
- **Tooling standards:** first-class MCP support, Agent Skills support and structured outputs save integration work.
- **Human-in-the-loop:** built-in approval pauses and resumable runs.
- **Observability:** tracing of every model call and tool call, exportable to your stack (many support OpenTelemetry).
- **State and durability:** can a run survive a crash, wait days for an approval and resume?
- **Security model:** sandboxing for code and files, secret handling, permission policies.
- **Team fit:** languages your team uses, maturity, documentation and community.

## Worked example: choosing for three teams

- **A two-person agency** automating lead research and CRM updates chooses **n8n** with its AI Agent node, MCP client tools and human review steps on the tools that write to the CRM. No servers to run beyond the n8n instance; business staff can read the flow.
- **A SaaS scale-up** building an in-product support agent chooses an **SDK loop helper** on the model provider's API, with its own tools, tracing and evaluation, because it needs tight control over latency, UX and data.
- **An enterprise data team** running nightly, multi-hour research jobs across internal documents chooses a **hosted agent platform** with sandboxed workspaces and scheduled runs, reviewed against its security requirements, so it does not operate long-running containers itself.

None of these is "best"; each matches the team's control, skills and risk needs.

## Hands-on: the same agent in two SDKs

**Anthropic Python SDK tool runner (beta helper that runs the loop for your tools):**

```python
import anthropic
from anthropic import beta_tool

client = anthropic.Anthropic()

@beta_tool
def get_exchange_rate(base: str, quote: str) -> str:
    """Return today's reference exchange rate between two ISO currency codes, e.g. AED and PKR."""
    rates = {("AED", "PKR"): 76.0, ("GBP", "AED"): 4.9}   # illustrative; call your real rates API
    rate = rates.get((base.upper(), quote.upper()))
    return f"{rate}" if rate else "Unknown pair; ask the user to confirm the currencies."

runner = client.beta.messages.tool_runner(
    model="claude-opus-5",            # check the current model list
    max_tokens=4000,
    tools=[get_exchange_rate],
    messages=[{"role": "user", "content": "Roughly how many PKR is AED 2,500? Show the rate you used."}],
)
for message in runner:
    final = message
print("".join(b.text for b in final.content if b.type == "text"))
```

**OpenAI Agents SDK (provider-agnostic agent framework):**

```python
from agents import Agent, Runner, function_tool

@function_tool
def get_exchange_rate(base: str, quote: str) -> str:
    """Return today's reference exchange rate between two ISO currency codes."""
    rates = {("AED", "PKR"): 76.0, ("GBP", "AED"): 4.9}   # illustrative values
    rate = rates.get((base.upper(), quote.upper()))
    return f"{rate}" if rate else "Unknown pair; ask the user to confirm the currencies."

agent = Agent(name="FX helper", instructions="Answer briefly and show the rate used.",
              tools=[get_exchange_rate])
result = Runner.run_sync(agent, "Roughly how many PKR is AED 2,500?")
print(result.final_output)
```

Install with `pip install anthropic` and `pip install openai-agents`, set `ANTHROPIC_API_KEY` or `OPENAI_API_KEY`, and run both. Notice how similar they are: a typed function with a docstring becomes a tool; the framework runs the loop. The differences that matter show up later, in tracing, approvals, state, sandboxes and model choice.

## Avoiding framework regret

- Keep **business logic in plain functions** and tool definitions you own; frameworks should call them, not contain them.
- Keep **prompts, tool descriptions and evaluation sets** in version control, framework-independent.
- Wrap model calls behind a thin interface so you can switch providers for a route.
- Run your **evaluation suite** on any candidate before migrating.
- The flagship **Building Production AI Agents** course builds production agents on these foundations; **Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API** covers provider APIs in depth.

## Go deeper

Related courses: for coding agents specifically, see **AI-Assisted Software Development: Coding Agents in Practice**; to run open-weight models yourself, see **Open-Weight and Local AI: Run, Choose and Deploy Your Own Models**; for agents that operate browsers and desktop apps, see **Computer-Use and Browser Agents**.

## Video lecture: Agent frameworks, SDKs and platforms: how to choose

Lecture coming soon · 16 chapters · about 9 minutes. Read the full transcript below.

1. Choosing agent tooling
2. Analogy: getting a kitchen
3. Two questions
4. Six categories
5. Eight criteria
6. Simple example: a leave-booking agent
7. Three teams, three answers
8. Business example (illustrative)
9. Hands-on in the lesson
10. Avoid regret
11. Run a two-week spike
12. Common mistakes
13. Did you choose well?
14. Watch me do it: two SDKs
15. Recap
16. Try this now (30 minutes)

## Lecture transcript

### Choosing agent tooling

New agent frameworks launch almost every week, each promising to be the one you need. If you choose by hype, you'll rebuild within a year. In this lesson you'll learn two questions that cut through the noise, a map of the main categories with current examples, eight selection criteria, and a hands-on comparison where you build the same tiny agent in two different SDKs.

### Analogy: getting a kitchen

Here's an analogy. Choosing an agent framework is like choosing how to get a kitchen. You can build every cabinet yourself, buy flat-pack units and assemble them, hire a fitted-kitchen company, or rent a flat with a kitchen included. None is best. It depends on your skills, budget, how custom you need it and who fixes it when a hinge breaks. Frameworks are the same trade-off between control and convenience.

### Two questions

Question one: who supplies the harness? That's the agent loop, context management, tool execution, retries, approvals and tracing. Question two: who supplies the deployment? Where the loop runs, where tools execute, and who keeps it running. If you write the loop against a raw API, you supply both. SDK helpers supply the loop, but you host it. Hosted agent platforms supply both, running the loop and a sandbox for you.

### Six categories

Here are six categories. Raw API plus your own loop, for maximum control. SDK loop helpers, like the Anthropic SDK's tool runner or the OpenAI Agents SDK, for custom tools on your infrastructure with less boilerplate. Full harness libraries, like the Claude Agent SDK, which packages Claude Code's harness with file and shell tools. Graph and workflow frameworks, like LangGraph or CrewAI, for explicit, auditable state machines. Hosted agent platforms from model providers and the big clouds. And no-code builders, like n8n's AI Agent node, Zapier Agents or Copilot Studio, for business teams.

### Eight criteria

Now the criteria. Control flow: do you need an explicit graph, a model-directed loop, or both? Model portability versus native features. Support for standards like MCP, Agent Skills and structured outputs. Built-in human approval and resumable runs. Observability, ideally exportable traces. Durability: can a run wait two days for an approval and resume? Security: sandboxing, secrets and permission policies. And team fit: languages, documentation and maturity. Score candidates against these, not against launch videos.

### Simple example: a leave-booking agent

A simple example of the two questions in action. You want a small agent that answers staff questions about the holiday policy and books leave in the HR system. If you have one developer and want full control, an SDK helper on your model provider's API gives you the loop without boilerplate, and you host it. If you have no developers, a no-code builder with an HR connector and an approval step does it without code, and the platform hosts it.

### Three teams, three answers

Here's how three teams chose. A two-person agency automating lead research picks n8n, with the AI Agent node, MCP client tools and a human review step on anything that writes to the CRM. A SaaS scale-up building an in-product support agent picks an SDK loop helper, because it needs tight control over latency, UX and data. An enterprise data team running multi-hour research jobs every night picks a hosted platform with sandboxed workspaces and scheduled runs, after a security review. Different answers, each right for its team.

### Business example (illustrative)

A deeper business example, illustrative. The SaaS scale-up first prototyped its support agent in a no-code builder, which proved the idea in a week. But it needed sub-second first responses, in-product context and custom tracing, so it rebuilt on an SDK helper. Because the tool functions, prompts and evaluation set were already separate files, the rebuild took about two weeks, and the evaluation suite showed identical quality on day one.

### Hands-on in the lesson

In the hands-on section you'll build the same small agent twice: a currency helper with one exchange-rate tool. Once with the Anthropic SDK's tool runner, once with the OpenAI Agents SDK. You'll see how alike they are: a typed function with a docstring becomes a tool, and the framework runs the loop. The differences show up later, in tracing, approvals, state, sandboxes and which models you can use.

### Avoid regret

Finally, avoid framework regret. Keep business logic in plain functions you own, and let frameworks call them. Keep prompts, tool descriptions and evaluation sets in version control, independent of any framework. Put a thin interface around model calls so you can switch providers for a route. And run your evaluation suite on any candidate before migrating. Products do change: some visual builders have already been retired, while their code-based SDKs continue.

### Run a two-week spike

One more practical tip before you commit: run a two-week spike. Take your evaluation suite, a realistic slice of traffic, and the two strongest candidates. Build the same agent in both, and compare success rates, cost per completed task, latency, how easy traces are to read, and how approvals feel to the people who'll use them. Two weeks of evidence beats two months of debate, and the spike leaves you with working code either way.

### Common mistakes

Common mistakes. Choosing a framework before you've written a single evaluation case. Picking the most popular option without checking it supports your model provider or MCP. Letting a framework's abstractions hide the actual prompts and tool calls, so you can't debug. Building on a visual builder for something that must run for years, without checking its roadmap. And migrating everything at once instead of one route at a time.

### Did you choose well?

How will you know you chose well? Six months in, your team can debug a bad run in minutes by reading traces. Adding a new tool takes an afternoon, not a sprint. Switching a route to a different model is a configuration change plus an evaluation run. And your core assets, prompts, tools and evaluations, would survive a move to another framework. If any of these are painful, reassess before you build more on top.

### Watch me do it: two SDKs

Watch me do it. First, the Anthropic version. I decorate get exchange rate with beta tool. The type hints for base and quote become the schema, and the docstring becomes the description. Inside, I look up an illustrative rate and return a helpful message for unknown pairs. Next, I create the tool runner with a model, max tokens, the tool and the user's question, then iterate over the runner until it finishes, keeping the final message, and print its text. Now the OpenAI Agents SDK version. I decorate the same function with function tool, create an Agent with a name, instructions and the tool, and call Runner run sync with the question. I print final output. Both answer: about a hundred and ninety thousand rupees, at a rate of seventy-six. Finally, I open each SDK's trace or log and check that I can see the tool call and its arguments, because that's what I'll rely on when things go wrong.

### Recap

To recap: ask who supplies the harness and who supplies the deployment, map options to six categories, score them on eight criteria, and keep your core assets portable. Your next step is to score two candidates for one agent you plan and build the hello-world in one of them. When you're ready to build production agents, the Building Production AI Agents course takes you much further.

### Try this now (30 minutes)

Try this now. Take one agent you plan to build. Answer the two questions, harness and deployment, in one sentence each. Score your top two candidates on the eight criteria from this lesson, using a simple one-to-five scale. Then build the currency hello-world from the hands-on section in the winner, and look at what its trace or log shows you. If you can't see the tool calls clearly, that's an important finding.

## Key takeaways

- Separate options by who supplies the harness (loop, context, tools) and who supplies the deployment.
- Categories: raw API loops, SDK helpers, full harness libraries, graph frameworks, hosted platforms and no-code builders.
- Choose on control flow, portability, MCP and Skills support, human-in-the-loop, observability, durability, security and team fit.
- Keep business logic, prompts and evaluations framework-independent to avoid lock-in and regret.

## Try it

Score two candidate frameworks or platforms for one agent you plan, using the eight selection criteria. Build the FX hello-world in one of them and note what tracing and approvals it offers.

- [Previous: Context engineering: curating what the model sees](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/context-engineering-fundamentals)
- [Next: MCP concepts: hosts, clients, servers and primitives](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/mcp-concepts)
- [All lessons of Latest AI Techniques: RAG, Tool Use, Agents & MCP](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp)
