Building Production AI AgentsTools, reasoning and multi-agent design · Lesson 6 of 18

Multi-agent systems: orchestration, handoffs and delegation

Article · 16 min · 9 min lecture

Video lecture

Multi-agent systems: orchestration, handoffs and delegation

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Multi-agent systems

  • When one agent isn't enough
  • Four topologies
  • Writing delegation briefs

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

When one agent is not enough

A single agent with good tools handles most tasks. Multi-agent designs help in three situations:

  1. Breadth: the work fans out into many independent investigations (research across 20 competitors).
  2. Context isolation: sub-tasks would flood one context window with reading; a sub-agent reads the 50 pages and returns a one-page summary.
  3. Specialization: different parts need different instructions, tools, permissions or models (a billing specialist that alone can issue refunds).

The price: multi-agent systems use many more tokens than a single agent, add coordination bugs, and are harder to evaluate. Anthropic has publicly described its multi-agent research system as outperforming a single agent on broad research tasks while consuming several times more tokens, so reserve these designs for high-value work.

Four topologies

TopologyHow it worksGood for
Orchestrator + sub-agentsLead agent plans, spawns sub-agents with focused briefs, synthesizes resultsResearch, audits, large codebases
HandoffsControl passes from one agent to a specialist that continues the conversationCustomer service triage (sales → billing → tech)
Agents as toolsA specialist agent is exposed as a tool; the caller keeps controlReusable skills like "translate and localize"
Pipelines / graphsFixed or conditional graph of agent nodesContent production, approvals, compliance flows

Frameworks map onto these: the OpenAI Agents SDK has first-class handoffs and agents as tools; LangGraph builds graphs with shared state; the Claude Agent SDK and Claude Code support sub-agents with their own context and tool permissions; CrewAI models crews of role-based agents plus flows; Google's ADK composes hierarchies and workflow agents; Microsoft Agent Framework (the successor to AutoGen and Semantic Kernel) offers multi-agent orchestration patterns. The concepts transfer; the APIs differ.

Writing good delegation briefs

Sub-agents fail when their brief is vague. Each brief should state:

  • Objective: one specific question.
  • Output format: e.g., JSON with findings, sources, confidence.
  • Boundaries: which tools and sources are allowed; what is out of scope.
  • Effort budget: e.g., "at most 8 tool calls".
  • Stop condition: "stop when you have three independent sources or after 8 calls".

Without boundaries, parallel sub-agents duplicate work or wander.

Worked example: competitor pricing audit for a Riyadh SaaS company

Goal: "Summarize pricing and packaging for our 12 main competitors in KSA and the GCC."

  1. Lead agent (stronger model, high effort) lists competitors and creates 12 briefs.
  2. Twelve sub-agents (cheaper model, low-to-medium effort) each research one competitor with web search and fetch tools, returning structured JSON with sources.
  3. Lead agent validates JSON, spots gaps (two competitors hide pricing), and spawns one follow-up sub-agent to check public procurement listings or reseller pages.
  4. Lead synthesizes a comparison table and flags low-confidence cells.

Cost control: sub-agents run in parallel with a concurrency limit, each capped at eight tool calls; the lead only reads summaries, never raw pages.

Hands-on: orchestrator with sub-agents in plain Python

import os, json, asyncio
import anthropic

client = anthropic.AsyncAnthropic()
LEAD = os.environ.get("LEAD_MODEL", "claude-opus-5")      # verify IDs in current docs
WORKER = os.environ.get("WORKER_MODEL", "claude-haiku-4-5")
SEM = asyncio.Semaphore(4)                                # stay inside rate limits

BRIEF = """Objective: {q}
Return JSON only: {{"findings": [..], "sources": [..], "confidence": "low|medium|high"}}
Boundaries: public information only; no guessing prices. Stop after you have 3 sources."""

async def worker(question: str) -> dict:
    async with SEM:
        msg = await client.messages.create(
            model=WORKER, max_tokens=2000,
            messages=[{"role": "user", "content": BRIEF.format(q=question)}],
            # add web search / fetch tools here; see the provider docs for tool types
        )
    text = "".join(b.text for b in msg.content if b.type == "text")
    try:
        return json.loads(text[text.find("{"): text.rfind("}") + 1])
    except json.JSONDecodeError:
        return {"findings": [], "sources": [], "confidence": "low", "error": "unparseable"}

async def orchestrate(goal: str, subquestions: list[str]) -> str:
    results = await asyncio.gather(*(worker(q) for q in subquestions))
    summary = json.dumps(dict(zip(subquestions, results)), ensure_ascii=False)[:60000]
    final = await client.messages.create(
        model=LEAD, max_tokens=4000,
        messages=[{"role": "user", "content": f"Goal: {goal}\nSub-agent results:\n{summary}\n"
                   "Synthesize a comparison. Flag low-confidence items and gaps."}],
    )
    return "".join(b.text for b in final.content if b.type == "text")

In a full system the lead would generate subquestions itself (structured output), and each worker would run the tool loop from lesson 3.

Agent-to-agent across organizations

Inside one codebase, sub-agents are just function calls. Across vendors and companies, open protocols are emerging: MCP connects agents to tools and data, while the Agent2Agent (A2A) protocol, now hosted under the Linux Foundation, targets communication between independent agents. Treat cross-organization agent traffic like any external API: authenticate, authorize, validate, log.

Pitfalls

  • Spawning sub-agents for tasks a single tool call could answer.
  • Sub-agents that return raw text dumps instead of structured summaries.
  • Shared mutable state without ownership rules; two agents overwrite each other.
  • Handoff loops (A hands to B, B hands back to A); cap handoffs per conversation.

Measuring success

Compare against a single-agent baseline on the same eval: quality, total tokens, wall-clock time and failure modes. Keep multi-agent only where the quality gain justifies the multiplier.

Key takeaways

  • Use multi-agent designs for breadth, context isolation or specialization, not by default.
  • Expect a large token multiplier; reserve multi-agent systems for high-value tasks.
  • Four topologies: orchestrator with sub-agents, handoffs, agents as tools, and graphs.
  • Delegation briefs need an objective, output format, boundaries, budget and stop condition.
  • Always benchmark against a single-agent baseline on the same eval.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A support bot must pass conversations to a billing specialist who continues talking to the customer. Which topology fits?
  2. What is the main benefit of giving 50-page reading tasks to sub-agents?
  3. Parallel sub-agents keep researching the same sources. What is the best fix?

Put it into practice

Take a research task from your work, write three sub-agent briefs using the five-part template, and estimate tokens for single-agent versus multi-agent versions.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.