Latest AI Techniques: RAG, Tool Use, Agents & MCPAgent loops, multi-agent patterns and memory · Lesson 10 of 20
Multi-agent patterns
Video lecture
Multi-agent patterns
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Multi-agent patterns
If one agent is good, surely five agents are better? Sometimes. Often not. Multi-agent systems can research faster, stay more focused and apply separation of duties, but they also multiply cost and create new ways to fail. In this lesson you'll learn when splitting work across agents pays off, five common patterns, how to write hand-offs that actually work, and how to stop errors from being laundered into facts.
0:30 Analogy: the restaurant kitchen
Here's an analogy. A single agent doing everything is like one person running a restaurant alone: taking orders, cooking, serving and washing up. It works for a tiny café. A busy restaurant has a head chef, line cooks and a waiter, each focused, working in parallel. But notice: the restaurant also needs clear tickets passing between them, or orders get lost. Multi-agent systems live or die by those tickets, the hand-offs.
1:01 Benefits vs costs
Why more than one? A single agent with many tools and a long transcript gets confused. Its context fills with details from unrelated sub-tasks, and one prompt has to cover many roles. Splitting gives you focus, because each agent sees only what it needs; parallelism, because independent sub-tasks run at once; and separation of duties, like a reviewer who didn't write the draft. The costs are more tokens, coordination latency and miscommunication. So start with the simplest design, and add agents only when a single agent demonstrably struggles.
1:39 Five patterns
Five patterns cover most designs. Orchestrator and workers: a lead agent breaks the task down, delegates to workers with fresh context, and synthesises. A pipeline of specialists: researcher, then writer, then editor, really a workflow whose steps are agents. Generator and critic: one produces, another evaluates against criteria, for a bounded number of rounds. Router and specialists: triage sends billing, technical and sales requests to agents with their own tools and permissions. And debate or voting: several agents answer independently and a judge or majority decides, at multiplied cost.
2:18 Write real hand-offs
Hand-offs are where multi-agent systems break. Research competitor B is a terrible delegation. Workers duplicate effort, go too deep or too shallow, and return things the orchestrator can't combine. A good brief has an objective, a definition of done, the exact output format expected back, constraints such as do not guess prices that aren't public, and only the context that worker needs. With structured outputs, the orchestrator can validate each worker's result against a schema before using it.
2:52 Simple example: 3 refund policies
A simple example. You need a short comparison of three competitors' refund policies. An orchestrator sends three workers one competitor each, with the same brief: summarise the refund window, conditions and exceptions, with a source link, and record anything not public as a gap. The workers run in parallel, return structured findings, and the orchestrator builds one comparison table. Three times faster than one agent reading every site in sequence, and each worker's context stays small.
3:25 Worked example: market-entry brief
Here's a worked example. A consultancy prepares market-entry briefs for e-commerce clients looking at Saudi Arabia or the UK. The orchestrator plans five questions: demand signals, regulation, payment preferences, logistics and competitors. Five workers run in parallel with search tools, each returning structured findings with sources. A critic checks every finding for support and flags anything unsupported or stale. The orchestrator writes the brief, and a human consultant reviews it before the client sees it. Faster and better sourced, at a higher token cost they judge worthwhile.
4:03 Business example (illustrative)
More on the market-entry case, illustrative. Sequentially, one agent took about forty minutes per brief and often ran out of useful context by the fourth question. In parallel, five workers finished in around ten minutes, each with a small, focused context. The critic flagged roughly one finding in eight as unsupported or out of date, which is exactly the kind of error that would otherwise have reached a client. Token cost roughly doubled, which the consultancy judged worth it.
4:37 Watch for
Now the dangers. Over-decomposition, ten agents where two would do. Duplicate work from vague boundaries. Lost context, where a worker misses a constraint the orchestrator knew. Runaway spawning, when agents can create agents without limits, so cap depth and count. And the sneakiest: error laundering. One worker's unsupported claim becomes accepted fact in the final synthesis, because it looks authoritative by the time it arrives. The cure is to carry sources through every hand-off and verify before synthesis. Least privilege per agent helps too: the researcher browses but can't send email.
5:17 Hands-on in the lesson
The hands-on section shows an orchestrator that fans out to parallel workers using the async Claude client, with each worker returning a validated structured finding, including sources and gaps. It caps the fan-out, tolerates worker failures and reports them honestly. You'll also meet the difference between agents as tools, where the orchestrator stays in charge, and hand-offs, where a specialist takes over the conversation, as in triage.
5:46 Common mistakes
Common mistakes, beyond the dangers you've heard. Splitting a task that isn't actually parallel, so workers wait on each other. Giving every worker every tool. Letting workers see the entire history, which defeats the point of fresh context. Using the most expensive model for simple reading tasks. And evaluating only the final output, so you never learn which agent in the chain is weak.
6:14 Does it pay off?
How will you know the multi-agent design pays off? Compare it with a single agent on the same tasks. Quality should be higher or time should be meaningfully lower, at a token cost you've decided is worth it. Every claim in the final output should trace to a worker's source. And each role should pass its own checks: sensible plans from the orchestrator, complete findings from workers, caught errors from the critic.
6:45 Watch me do it: async orchestrator
Watch me do it. I open the orchestrator. First, the Finding model: question, answer, sources and gaps. Next, worker calls messages dot parse on the async client with a system prompt that says use only the provided material and record unknowns as gaps, and returns a validated Finding. Then orchestrate builds up to five briefs, each with a task, a done-when rule, constraints and only the relevant material. Asyncio gather runs them in parallel, with return exceptions true so one failure doesn't sink the rest. I split results into findings and failed questions. Finally, a synthesis call writes the one-page brief, keeping every source and listing gaps and failures honestly. I run it on five market questions. The console shows five workers start at once, four findings return, one fails on a timeout, and the brief ends with a section called gaps and failed questions.
7:48 Recap
To recap: multi-agent designs buy focus, parallelism and separation of duties at the price of tokens and coordination. Use the five patterns deliberately, write structured hand-offs, apply least privilege per agent, cap spawning, and carry sources through to prevent error laundering. Your next step is to sketch a two- or three-agent design for a task you know and write one worker's delegation brief. Next, we'll give agents memory.
8:18 Try this now (20 minutes)
Try this now. Sketch a two- or three-agent design for a task you know, such as a weekly competitor digest or a client onboarding pack. Decide the pattern: orchestrator and workers, a specialist pipeline, or generator and critic. Then write one worker's delegation brief with the task, the definition of done, the output format, constraints and the sources it must return. Ask a colleague whether they could do the job from that brief alone.
Why more than one agent?
A single agent with many tools and a long transcript can become confused: its context fills with details from unrelated sub-tasks, and one prompt must cover many roles. Multi-agent designs split work across several model instances, each with a focused role, prompt, tools and context.
The benefits: focus (each agent sees only what it needs), parallelism (independent sub-tasks run at once), and separation of duties (a reviewer that did not write the draft). The costs: more tokens, more latency in coordination, and more places for miscommunication. Many experienced practitioners recommend starting with the simplest design that works and adding agents only when a single agent demonstrably struggles.
Common patterns
1. Orchestrator and workers (supervisor). A lead agent breaks the task into sub-tasks, delegates each to a worker agent with fresh context, and synthesises results. Good for broad research: each worker investigates one competitor, one market or one question in parallel.
2. Pipeline of specialists. Fixed stages, each handled by a specialised agent: researcher, then writer, then editor. Really a workflow whose steps are agents.
3. Generator and critic. One agent produces, another evaluates against criteria and returns feedback, looping a bounded number of times. The critic benefits from different instructions and, ideally, different information (for example the source documents).
4. Router and specialists. A triage step routes requests to specialist agents (billing, technical, sales), each with its own tools and permissions.
5. Debate or voting. Several agents answer independently; a judge or majority vote selects. Can improve reliability on hard judgement questions, at multiplied cost.
Designing hand-offs
Hand-offs are where multi-agent systems fail. When delegating, the orchestrator should give each worker:
- a clear objective and the definition of done
- the output format expected back
- boundaries (what not to do, which tools to use)
- relevant context, but not the entire history
{
"task": "Summarise publicly available pricing for Competitor B",
"done_when": "Tiers, prices, billing period and main limits listed, each with a source URL",
"output_format": {"tiers": [{"name": "", "price": "", "period": "", "limits": "", "source": ""}], "gaps": []},
"constraints": ["Do not guess prices not shown publicly; record them in gaps"]
}Vague delegation ("research competitor B") produces workers that duplicate effort, go too deep or too shallow, and return results the orchestrator cannot combine.
Permissions per agent
Multi-agent designs let you apply least privilege: the researcher can browse but not send email; the drafter can write drafts but not publish; only a step with human approval can take consequential actions. This is a real security advantage when well designed.
Worked example: a market entry brief
A consultancy prepares briefs on entering new markets (for example Saudi Arabia or the UK) for e-commerce clients.
- Orchestrator plans five research questions: market size signals, regulation, payment preferences, logistics, competitors.
- Five workers run in parallel with web search tools, each returning structured findings with sources.
- A critic agent checks each finding for source support and flags anything unsupported or outdated.
- The orchestrator writes the brief; a human consultant reviews before it goes to the client.
Compared with one agent doing everything sequentially, the brief is completed faster and each section is better sourced, at a higher token cost that the consultancy judges worthwhile for this task.
Failure modes
- Over-decomposition: ten agents where two would do, multiplying cost and coordination errors.
- Duplicate work: overlapping sub-tasks because boundaries were vague.
- Lost context: a worker lacks a key constraint the orchestrator knew.
- Error laundering: a worker's unsupported claim becomes "fact" in the final synthesis. Carry sources through and verify.
- Runaway spawning: agents that can spawn agents without limits. Cap depth and count.
Hands-on: orchestrator and parallel workers with structured hand-offs
The sketch below uses the Claude API with Python's asyncio to run workers in parallel. Each worker receives a structured delegation brief and must return JSON matching a schema, which the orchestrator validates before synthesis. Structured outputs (covered in module 3) make the hand-off reliable.
import asyncio, json, os
import anthropic
from pydantic import BaseModel
client = anthropic.AsyncAnthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
class Finding(BaseModel):
question: str
answer: str
sources: list[str]
gaps: list[str]
async def worker(brief: dict) -> Finding:
resp = await client.messages.parse(
model=MODEL, max_tokens=4000,
system="You are a careful research worker. Use only the provided material. "
"Never invent figures; record unknowns in gaps.",
messages=[{"role": "user", "content": json.dumps(brief)}],
output_format=Finding,
)
return resp.parsed_output
async def orchestrate(topic: str, questions: list[str], material: str):
briefs = [{"task": q, "done_when": "answered with at least one source, or gap recorded",
"constraints": ["no guessing", "cite the material section"], "material": material}
for q in questions[:5]] # cap the fan-out
results = await asyncio.gather(*(worker(b) for b in briefs), return_exceptions=True)
findings = [r for r in results if isinstance(r, Finding)]
failed = [q for q, r in zip(questions, results) if not isinstance(r, Finding)]
synthesis = await client.messages.create(
model=MODEL, max_tokens=4000,
messages=[{"role": "user", "content":
f"Write a one-page brief on {topic} from these findings. Keep every source; "
f"list gaps and failed questions honestly.\n{[f.model_dump() for f in findings]}\n"
f"Failed: {failed}"}])
return "".join(b.text for b in synthesis.content if b.type == "text")Production notes: pass each worker only the context it needs (here, the relevant material section rather than everything); set a cap on fan-out and depth; give each worker the minimum tools; use a cheaper, faster model for reading-heavy workers when evaluation shows quality holds; and log every brief and result under one run ID so any claim in the final brief can be traced to its worker and source.
Agents as tools versus hand-offs
Frameworks describe two ways agents cooperate. With agents as tools, the orchestrator calls a specialist like a function and keeps control. With hand-offs, control transfers to the specialist, which then talks to the user (common in triage: "billing agent, take over"). The OpenAI Agents SDK, for example, supports both explicitly, and hosted agent platforms increasingly offer multi-agent sessions where an agent can delegate to copies of itself or to named worker agents. Pick agents-as-tools when one voice should own the final answer; pick hand-offs when the specialist should own the conversation.
Going further
Evaluate multi-agent systems end to end and per role: did the orchestrator plan sensible sub-tasks? Did workers meet their definitions of done? Did the critic catch planted errors? Log each agent's inputs and outputs with a shared run ID so you can trace any final claim back to its origin.
Key takeaways
- Multi-agent designs offer focus, parallelism and separation of duties, at the cost of tokens and coordination.
- Patterns: orchestrator-workers, specialist pipeline, generator-critic, router-specialists, debate or voting.
- Hand-offs need objectives, definitions of done, output formats, boundaries and relevant context.
- Apply least privilege per agent, cap spawning, and carry sources through to prevent error laundering.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Sketch a two- or three-agent design for a task you know. Write the delegation message for one worker using the JSON structure above.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.