---
title: "Multi-agent systems: orchestration, handoffs and delegation"
description: "When one agent is not enough A single agent with good tools handles most tasks. Multi-agent designs help in three situations: 1. Breadth : the work fans…"
url: https://optimizeall.com/learn/ai-agents-engineering/multi-agent-systems
updated: 2026-10-05
---

Building Production AI Agents · Tools, reasoning and multi-agent design · lesson 6 of 18 · 16 min

# Multi-agent systems: orchestration, handoffs and delegation

## When one agent is not enough

A single agent with good tools handles most tasks. Multi-agent designs help in three situations:

1. **Breadth**: the work fans out into many independent investigations (research across 20 competitors).
2. **Context isolation**: sub-tasks would flood one context window with reading; a sub-agent reads the 50 pages and returns a one-page summary.
3. **Specialization**: different parts need different instructions, tools, permissions or models (a billing specialist that alone can issue refunds).

The price: multi-agent systems use many more tokens than a single agent, add coordination bugs, and are harder to evaluate. Anthropic has publicly described its multi-agent research system as outperforming a single agent on broad research tasks while consuming several times more tokens, so reserve these designs for high-value work.

## Four topologies

| Topology | How it works | Good for |
|---|---|---|
| **Orchestrator + sub-agents** | Lead agent plans, spawns sub-agents with focused briefs, synthesizes results | Research, audits, large codebases |
| **Handoffs** | Control passes from one agent to a specialist that continues the conversation | Customer service triage (sales → billing → tech) |
| **Agents as tools** | A specialist agent is exposed as a tool; the caller keeps control | Reusable skills like "translate and localize" |
| **Pipelines / graphs** | Fixed or conditional graph of agent nodes | Content production, approvals, compliance flows |

Frameworks map onto these: the OpenAI Agents SDK has first-class **handoffs** and **agents as tools**; LangGraph builds **graphs** with shared state; the Claude Agent SDK and Claude Code support **sub-agents** with their own context and tool permissions; CrewAI models **crews** of role-based agents plus **flows**; Google's ADK composes hierarchies and workflow agents; Microsoft Agent Framework (the successor to AutoGen and Semantic Kernel) offers multi-agent orchestration patterns. The concepts transfer; the APIs differ.

## Writing good delegation briefs

Sub-agents fail when their brief is vague. Each brief should state:

- **Objective**: one specific question.
- **Output format**: e.g., JSON with `findings`, `sources`, `confidence`.
- **Boundaries**: which tools and sources are allowed; what is out of scope.
- **Effort budget**: e.g., "at most 8 tool calls".
- **Stop condition**: "stop when you have three independent sources or after 8 calls".

Without boundaries, parallel sub-agents duplicate work or wander.

## Worked example: competitor pricing audit for a Riyadh SaaS company

Goal: "Summarize pricing and packaging for our 12 main competitors in KSA and the GCC."

1. Lead agent (stronger model, high effort) lists competitors and creates 12 briefs.
2. Twelve sub-agents (cheaper model, low-to-medium effort) each research one competitor with web search and fetch tools, returning structured JSON with sources.
3. Lead agent validates JSON, spots gaps (two competitors hide pricing), and spawns one follow-up sub-agent to check public procurement listings or reseller pages.
4. Lead synthesizes a comparison table and flags low-confidence cells.

Cost control: sub-agents run in parallel with a concurrency limit, each capped at eight tool calls; the lead only reads summaries, never raw pages.

## Hands-on: orchestrator with sub-agents in plain Python

```python
import os, json, asyncio
import anthropic

client = anthropic.AsyncAnthropic()
LEAD = os.environ.get("LEAD_MODEL", "claude-opus-5")      # verify IDs in current docs
WORKER = os.environ.get("WORKER_MODEL", "claude-haiku-4-5")
SEM = asyncio.Semaphore(4)                                # stay inside rate limits

BRIEF = """Objective: {q}
Return JSON only: {{"findings": [..], "sources": [..], "confidence": "low|medium|high"}}
Boundaries: public information only; no guessing prices. Stop after you have 3 sources."""

async def worker(question: str) -> dict:
    async with SEM:
        msg = await client.messages.create(
            model=WORKER, max_tokens=2000,
            messages=[{"role": "user", "content": BRIEF.format(q=question)}],
            # add web search / fetch tools here; see the provider docs for tool types
        )
    text = "".join(b.text for b in msg.content if b.type == "text")
    try:
        return json.loads(text[text.find("{"): text.rfind("}") + 1])
    except json.JSONDecodeError:
        return {"findings": [], "sources": [], "confidence": "low", "error": "unparseable"}

async def orchestrate(goal: str, subquestions: list[str]) -> str:
    results = await asyncio.gather(*(worker(q) for q in subquestions))
    summary = json.dumps(dict(zip(subquestions, results)), ensure_ascii=False)[:60000]
    final = await client.messages.create(
        model=LEAD, max_tokens=4000,
        messages=[{"role": "user", "content": f"Goal: {goal}\nSub-agent results:\n{summary}\n"
                   "Synthesize a comparison. Flag low-confidence items and gaps."}],
    )
    return "".join(b.text for b in final.content if b.type == "text")
```

In a full system the lead would generate `subquestions` itself (structured output), and each worker would run the tool loop from lesson 3.

## Agent-to-agent across organizations

Inside one codebase, sub-agents are just function calls. Across vendors and companies, open protocols are emerging: **MCP** connects agents to tools and data, while the **Agent2Agent (A2A)** protocol, now hosted under the Linux Foundation, targets communication between independent agents. Treat cross-organization agent traffic like any external API: authenticate, authorize, validate, log.

## Pitfalls

- Spawning sub-agents for tasks a single tool call could answer.
- Sub-agents that return raw text dumps instead of structured summaries.
- Shared mutable state without ownership rules; two agents overwrite each other.
- Handoff loops (A hands to B, B hands back to A); cap handoffs per conversation.

## Measuring success

Compare against a single-agent baseline on the same eval: quality, total tokens, wall-clock time and failure modes. Keep multi-agent only where the quality gain justifies the multiplier.

## Video lecture: Multi-agent systems: orchestration, handoffs and delegation

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

1. Multi-agent systems
2. Why it matters
3. Why go multi-agent
4. Four topologies
5. The five-part brief
6. Simple example: five pricing pages
7. Example: 12-competitor pricing audit
8. Hands-on skeleton
9. Pitfalls
10. Clean hand-backs
11. Deeper: the Riyadh audit (illustrative)
12. Watch me do it: orchestrate()
13. Try this now
14. Recap

## Lecture transcript

### Multi-agent systems

Picture a research director with a team of analysts. She doesn't read every page herself. She hands out focused briefs, then combines what comes back. Multi agent AI systems work the same way, and when they fit, they're remarkable. But they're also expensive and tricky. In this lesson you'll learn when to use them, the four main topologies, and how to write briefs that sub agents actually follow.

### Why it matters

Why does this matter? Because multi agent demos are impressive, and that makes them tempting for everything. But every extra agent is another context to fill, another set of tokens to pay for, and another place for miscommunication. Think of a group project at school. Three people can research three topics faster than one person. But if the brief is vague, you get three overlapping essays and a stressful night stitching them together. The value of multi agent systems depends almost entirely on clear division of work and clean hand backs.

### Why go multi-agent

There are three good reasons to go multi agent. Breadth: the work fans out into many independent investigations, like researching twenty competitors. Context isolation: a sub agent can read fifty pages and return one page, keeping the lead's context clean. And specialization: some parts need different instructions, tools, permissions or even models, like a billing agent that alone can issue refunds. The cost is real, though. Multi agent systems burn many more tokens, and Anthropic has described its own research system as using several times the tokens of a single agent. So reserve it for valuable work.

### Four topologies

Now the four topologies. Orchestrator and sub agents: a lead plans, spawns focused workers and synthesizes. Handoffs: control passes to a specialist who continues the conversation, great for customer service triage. Agents as tools: a specialist is wrapped as a tool, and the caller stays in charge. And graphs or pipelines: agent nodes wired together with fixed or conditional edges. Frameworks lean toward different ones. The OpenAI Agents SDK has handoffs and agents as tools, LangGraph builds graphs, and the Claude Agent SDK supports sub agents with their own context and permissions.

### The five-part brief

Sub agents live or die by their brief. Every brief needs five parts. An objective: one specific question. An output format, such as JSON with findings, sources and confidence. Boundaries: which tools and sources are allowed, and what's out of scope. An effort budget, like eight tool calls at most. And a stop condition: stop when you have three independent sources. Leave out the boundaries and your parallel workers will duplicate each other or wander off.

### Simple example: five pricing pages

A simple example. You want a summary of five competitors' pricing pages. Single agent version: one agent reads all five pages, and its context fills with navigation menus and cookie banners. Multi agent version: the lead sends five sub agents one URL each, with the same brief, return plan names, prices, currency and a source link as JSON. Each sub agent reads one page and returns a few lines. The lead combines five tidy results. The total token count may be higher, but the lead's context stays clean, and the answer is easier to check.

### Example: 12-competitor pricing audit

A worked example. A software company in Riyadh wants pricing and packaging for twelve competitors across Saudi Arabia and the Gulf. The lead agent, on a stronger model, lists the competitors and writes twelve briefs. Twelve workers on a cheaper model research one competitor each and return structured JSON with sources. The lead notices two competitors hide their pricing, so it sends one follow up worker to check reseller and procurement pages. Finally it builds a comparison table and flags low confidence cells. Workers run four at a time, capped at eight tool calls each.

### Hands-on skeleton

The lesson's code shows the skeleton in plain Python. An async semaphore limits concurrency so you stay within rate limits. Workers use a smaller model and a strict brief. If a worker's JSON can't be parsed, it returns a low confidence placeholder instead of crashing everything. The lead only ever sees compact summaries, never raw pages. Across organizations, open protocols are emerging too. MCP connects agents to tools and data, and the Agent to Agent protocol, now under the Linux Foundation, targets agents talking to other agents. Treat that traffic like any external API: authenticate, validate and log.

### Pitfalls

Watch for four pitfalls. Spawning sub agents for questions a single tool call could answer. Workers that return raw dumps instead of structured summaries. Shared state with no ownership rules, so agents overwrite each other. And handoff loops, where A hands to B and B hands straight back, so cap handoffs per conversation. Most importantly, always benchmark against a single agent baseline on the same eval. Keep the multi agent version only if the quality gain is worth the multiplier.

### Clean hand-backs

How do sub agents hand results back without flooding the lead? Ask each one for a compact, structured summary: the findings in a few bullet points, the sources with links, a confidence level and any open questions. Cap the length, for example two hundred words. The lead then reads a handful of tidy summaries instead of dozens of raw pages. If the lead needs detail on one point, it can ask that specific sub agent a follow up question, rather than every sub agent sending everything up front.

### Deeper: the Riyadh audit (illustrative)

Let's deepen the Riyadh pricing audit with illustrative numbers. Twelve competitors, twelve sub agents, each capped at eight tool calls. The first run cost several times a single agent run, as expected. But it finished in a fraction of the time, because workers ran four at a time, and the comparison table was far more complete: the single agent had given up after seven competitors when its context filled with web pages. Two competitors hid their prices; the follow up worker found reseller listings for one and marked the other as not public. The sales team said the output replaced about two days of manual research per quarter, which easily justified the higher token bill.

### Watch me do it: orchestrate()

Watch me do it. Let's step through the orchestrator code. The lead model and worker model come from environment variables, so I can swap them without code changes. The semaphore is set to four, which caps concurrency. The brief template has the objective, a JSON output shape with findings, sources and confidence, and boundaries: public information only, no guessing prices, stop after three sources. The worker function acquires the semaphore, calls the worker model with the formatted brief, and tries to parse JSON from the reply. If parsing fails, it returns a low confidence placeholder instead of crashing the whole run. In orchestrate, asyncio gather runs all workers, then I zip questions with results into one JSON summary, truncated to a safe size, and send it to the lead model to synthesize, asking it to flag low confidence items and gaps.

### Try this now

Try this now. Take a research question you'd genuinely like answered, like how do three competitors in your market handle free trials. Write three sub agent briefs using the five parts: objective, output format, boundaries, budget and stop condition. Then run the question once with a single agent and once with the briefs, and compare quality, tokens and time. You might find the single agent is good enough, which is a perfectly valid result. The point is that you decided with evidence.

### Recap

To recap: use multi agent designs for breadth, isolation or specialization. Pick the right topology, write five part briefs, cap concurrency and calls, and prove the value against a single agent. Your next step: choose a research task from your own work, write three sub agent briefs, and estimate the token cost for single and multi agent versions.

## Key takeaways

- Use multi-agent designs for breadth, context isolation or specialization, not by default.
- Expect a large token multiplier; reserve multi-agent systems for high-value tasks.
- Four topologies: orchestrator with sub-agents, handoffs, agents as tools, and graphs.
- Delegation briefs need an objective, output format, boundaries, budget and stop condition.
- Always benchmark against a single-agent baseline on the same eval.

## Try it

Take a research task from your work, write three sub-agent briefs using the five-part template, and estimate tokens for single-agent versus multi-agent versions.

- [Previous: Planning, reasoning models and effort control](https://optimizeall.com/learn/ai-agents-engineering/planning-and-reasoning-models)
- [Next: Context engineering and agent memory](https://optimizeall.com/learn/ai-agents-engineering/context-engineering-and-memory)
- [All lessons of Building Production AI Agents](https://optimizeall.com/learn/ai-agents-engineering)
