Latest AI Techniques: RAG, Tool Use, Agents & MCPAgent loops, multi-agent patterns and memory · Lesson 10 of 20

Multi-agent patterns

Article · 11 min · 9 min lecture

Video lecture

Multi-agent patterns

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Multi-agent patterns

  • When splitting work pays off
  • Five patterns
  • Hand-offs and error laundering

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why more than one agent?

A single agent with many tools and a long transcript can become confused: its context fills with details from unrelated sub-tasks, and one prompt must cover many roles. Multi-agent designs split work across several model instances, each with a focused role, prompt, tools and context.

The benefits: focus (each agent sees only what it needs), parallelism (independent sub-tasks run at once), and separation of duties (a reviewer that did not write the draft). The costs: more tokens, more latency in coordination, and more places for miscommunication. Many experienced practitioners recommend starting with the simplest design that works and adding agents only when a single agent demonstrably struggles.

Common patterns

1. Orchestrator and workers (supervisor). A lead agent breaks the task into sub-tasks, delegates each to a worker agent with fresh context, and synthesises results. Good for broad research: each worker investigates one competitor, one market or one question in parallel.

2. Pipeline of specialists. Fixed stages, each handled by a specialised agent: researcher, then writer, then editor. Really a workflow whose steps are agents.

3. Generator and critic. One agent produces, another evaluates against criteria and returns feedback, looping a bounded number of times. The critic benefits from different instructions and, ideally, different information (for example the source documents).

4. Router and specialists. A triage step routes requests to specialist agents (billing, technical, sales), each with its own tools and permissions.

5. Debate or voting. Several agents answer independently; a judge or majority vote selects. Can improve reliability on hard judgement questions, at multiplied cost.

Designing hand-offs

Hand-offs are where multi-agent systems fail. When delegating, the orchestrator should give each worker:

  • a clear objective and the definition of done
  • the output format expected back
  • boundaries (what not to do, which tools to use)
  • relevant context, but not the entire history
{
  "task": "Summarise publicly available pricing for Competitor B",
  "done_when": "Tiers, prices, billing period and main limits listed, each with a source URL",
  "output_format": {"tiers": [{"name": "", "price": "", "period": "", "limits": "", "source": ""}], "gaps": []},
  "constraints": ["Do not guess prices not shown publicly; record them in gaps"]
}

Vague delegation ("research competitor B") produces workers that duplicate effort, go too deep or too shallow, and return results the orchestrator cannot combine.

Permissions per agent

Multi-agent designs let you apply least privilege: the researcher can browse but not send email; the drafter can write drafts but not publish; only a step with human approval can take consequential actions. This is a real security advantage when well designed.

Worked example: a market entry brief

A consultancy prepares briefs on entering new markets (for example Saudi Arabia or the UK) for e-commerce clients.

  • Orchestrator plans five research questions: market size signals, regulation, payment preferences, logistics, competitors.
  • Five workers run in parallel with web search tools, each returning structured findings with sources.
  • A critic agent checks each finding for source support and flags anything unsupported or outdated.
  • The orchestrator writes the brief; a human consultant reviews before it goes to the client.

Compared with one agent doing everything sequentially, the brief is completed faster and each section is better sourced, at a higher token cost that the consultancy judges worthwhile for this task.

Failure modes

  • Over-decomposition: ten agents where two would do, multiplying cost and coordination errors.
  • Duplicate work: overlapping sub-tasks because boundaries were vague.
  • Lost context: a worker lacks a key constraint the orchestrator knew.
  • Error laundering: a worker's unsupported claim becomes "fact" in the final synthesis. Carry sources through and verify.
  • Runaway spawning: agents that can spawn agents without limits. Cap depth and count.

Hands-on: orchestrator and parallel workers with structured hand-offs

The sketch below uses the Claude API with Python's asyncio to run workers in parallel. Each worker receives a structured delegation brief and must return JSON matching a schema, which the orchestrator validates before synthesis. Structured outputs (covered in module 3) make the hand-off reliable.

import asyncio, json, os
import anthropic
from pydantic import BaseModel

client = anthropic.AsyncAnthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")

class Finding(BaseModel):
    question: str
    answer: str
    sources: list[str]
    gaps: list[str]

async def worker(brief: dict) -> Finding:
    resp = await client.messages.parse(
        model=MODEL, max_tokens=4000,
        system="You are a careful research worker. Use only the provided material. "
               "Never invent figures; record unknowns in gaps.",
        messages=[{"role": "user", "content": json.dumps(brief)}],
        output_format=Finding,
    )
    return resp.parsed_output

async def orchestrate(topic: str, questions: list[str], material: str):
    briefs = [{"task": q, "done_when": "answered with at least one source, or gap recorded",
               "constraints": ["no guessing", "cite the material section"], "material": material}
              for q in questions[:5]]                       # cap the fan-out
    results = await asyncio.gather(*(worker(b) for b in briefs), return_exceptions=True)
    findings = [r for r in results if isinstance(r, Finding)]
    failed = [q for q, r in zip(questions, results) if not isinstance(r, Finding)]
    synthesis = await client.messages.create(
        model=MODEL, max_tokens=4000,
        messages=[{"role": "user", "content":
                   f"Write a one-page brief on {topic} from these findings. Keep every source; "
                   f"list gaps and failed questions honestly.\n{[f.model_dump() for f in findings]}\n"
                   f"Failed: {failed}"}])
    return "".join(b.text for b in synthesis.content if b.type == "text")

Production notes: pass each worker only the context it needs (here, the relevant material section rather than everything); set a cap on fan-out and depth; give each worker the minimum tools; use a cheaper, faster model for reading-heavy workers when evaluation shows quality holds; and log every brief and result under one run ID so any claim in the final brief can be traced to its worker and source.

Agents as tools versus hand-offs

Frameworks describe two ways agents cooperate. With agents as tools, the orchestrator calls a specialist like a function and keeps control. With hand-offs, control transfers to the specialist, which then talks to the user (common in triage: "billing agent, take over"). The OpenAI Agents SDK, for example, supports both explicitly, and hosted agent platforms increasingly offer multi-agent sessions where an agent can delegate to copies of itself or to named worker agents. Pick agents-as-tools when one voice should own the final answer; pick hand-offs when the specialist should own the conversation.

Going further

Evaluate multi-agent systems end to end and per role: did the orchestrator plan sensible sub-tasks? Did workers meet their definitions of done? Did the critic catch planted errors? Log each agent's inputs and outputs with a shared run ID so you can trace any final claim back to its origin.

Key takeaways

  • Multi-agent designs offer focus, parallelism and separation of duties, at the cost of tokens and coordination.
  • Patterns: orchestrator-workers, specialist pipeline, generator-critic, router-specialists, debate or voting.
  • Hand-offs need objectives, definitions of done, output formats, boundaries and relevant context.
  • Apply least privilege per agent, cap spawning, and carry sources through to prevent error laundering.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. When is an orchestrator-workers pattern most useful?
  2. What is 'error laundering' in multi-agent systems?
  3. Which delegation message is most likely to produce combinable results?

Put it into practice

Sketch a two- or three-agent design for a task you know. Write the delegation message for one worker using the JSON structure above.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.