Building Production AI AgentsAgent foundations and architectures · Lesson 2 of 18

Workflow and agent architectures: the pattern catalog

Article · 16 min · 9 min lecture

Video lecture

Workflow and agent architectures: the pattern catalog

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Workflow and agent patterns

  • Seven reusable building blocks
  • When each one fits
  • Combining them

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why patterns matter

Most successful LLM systems are built from a handful of composable patterns rather than sprawling frameworks. Knowing them lets you pick the simplest design that works and explain it to stakeholders. The catalog below follows the widely used taxonomy popularized by Anthropic's "Building effective agents" guidance, extended with what teams have learned since.

Pattern 1: Prompt chaining

Break a task into fixed sequential steps, each an LLM call, with optional gates (code checks) between them.

  • Use when: the task decomposes cleanly (outline → draft → edit).
  • Gate example: after generating an outline, check it has 4–7 sections before drafting.
  • Trade-off: more latency, much higher reliability per step.

Pattern 2: Routing

Classify the input, then send it to a specialized prompt, tool set or model.

  • Use when: inputs fall into distinct categories with different handling (refund vs technical issue vs sales inquiry).
  • Bonus: route easy queries to a small, cheap model and hard ones to a frontier model.
  • Trade-off: misroutes; always measure classifier accuracy.

Pattern 3: Parallelization

Run LLM calls concurrently and aggregate.

  • Sectioning: split independent subtasks (analyze five competitor pages at once).
  • Voting: run the same task several times and aggregate for confidence (for example, three independent reviews of whether an ad violates policy).
  • Trade-off: cost scales with branches; latency drops.

Pattern 4: Orchestrator–workers

A central model breaks the task into subtasks dynamically, delegates them to worker calls (often cheaper models), and synthesizes the results. Unlike parallelization, the subtasks are not known in advance.

  • Use when: "Update our pricing mentions across the site" — you don't know which pages need changes until you look.

Pattern 5: Evaluator–optimizer

One call generates, another evaluates against explicit criteria and returns feedback; loop until the evaluator passes or a limit is hit.

  • Use when: clear evaluation criteria exist and iteration measurably improves output (translation nuance, ad copy meeting a brand checklist, code passing tests).
  • Tip: the evaluator should be as specific as a rubric, not "is this good?".

Pattern 6: The autonomous agent

A single model in a loop with tools, deciding its own steps until done. Simple to write, hard to make reliable. Invest in tool design, environment feedback (tests, validators) and stopping conditions.

Pattern 7: Multi-agent systems

Several agents with separate contexts and roles: a lead agent with sub-agents, handoffs between specialists, or agents exposed to one another as tools. Powerful for breadth-first research and context isolation, but multi-agent systems typically consume many more tokens than a single agent, so reserve them for high-value tasks (module 2 goes deeper).

Choosing: a decision table

SituationStart with
Known steps, quality mattersPrompt chaining with gates
Distinct input categoriesRouting
Independent subtasks or need for confidenceParallelization
Unknown subtasks discovered at runtimeOrchestrator–workers
Clear rubric and iteration helpsEvaluator–optimizer
Open-ended, many possible pathsSingle agent with good tools
Very broad research or context overflowMulti-agent (lead + sub-agents)

Worked example: a UK e-commerce brand's product-description pipeline

A London brand needs 400 product descriptions in brand voice, compliant with UK advertising rules (no unsupported claims like "clinically proven" without evidence).

  1. Routing: classify each product (apparel, skincare, accessories). Skincare gets a stricter prompt about health claims.
  2. Prompt chaining: extract attributes from the spec sheet → draft → gate (length 80–120 words, required keywords present).
  3. Evaluator–optimizer: a checker prompt scores against a rubric (voice, claims, readability) and returns fixes; max two revision rounds.
  4. Parallelization: process products concurrently within rate limits.

No agent needed. The result is predictable, cheap and auditable. The one place they did add an agent: a "research unusual ingredient" step that can search approved sources when the spec sheet mentions an unfamiliar component.

Hands-on: an evaluator–optimizer in Python

import os, json
import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from the environment
MODEL = os.environ.get("MODEL", "claude-sonnet-5")  # check current model IDs in the docs

RUBRIC = """Score 1-5 on each: brand_voice (warm, plain English), claims (no unsupported
health or performance claims), length (80-120 words). Return JSON:
{"scores": {...}, "pass": true|false, "fixes": ["..."]}. pass only if all scores >= 4."""

def ask(system: str, user: str) -> str:
    msg = client.messages.create(model=MODEL, max_tokens=2000, system=system,
                                 messages=[{"role": "user", "content": user}])
    return "".join(b.text for b in msg.content if b.type == "text")

def generate(spec: str, fixes: list[str] | None = None) -> str:
    extra = f"\nApply these fixes: {fixes}" if fixes else ""
    return ask("You write product descriptions for a UK skincare brand.",
               f"Spec sheet:\n{spec}{extra}\nWrite the description only.")

def evaluate(draft: str) -> dict:
    raw = ask("You are a strict copy reviewer. Output JSON only.", f"{RUBRIC}\n\nDraft:\n{draft}")
    return json.loads(raw[raw.find("{"): raw.rfind("}") + 1])

def pipeline(spec: str, max_rounds: int = 3) -> dict:
    draft, verdict = generate(spec), {}
    for round_no in range(max_rounds):
        verdict = evaluate(draft)
        if verdict.get("pass"):
            return {"draft": draft, "rounds": round_no + 1, "verdict": verdict}
        draft = generate(spec, verdict.get("fixes"))
    return {"draft": draft, "rounds": max_rounds, "verdict": verdict, "needs_human": True}

Notice the round cap and the needs_human flag: the loop always terminates, and failures route to a person. In production you would use structured outputs (covered later) instead of slicing JSON from text.

Pitfalls

  • Using an agent where routing plus chaining would do.
  • Evaluators with vague criteria, which approve everything or loop forever.
  • Frameworks that hide prompts; you must be able to see and version every prompt.
  • Forgetting that parallel branches multiply rate-limit usage.

Measuring success

For each pattern, track per-step pass rates (gates, evaluator verdicts), end-to-end acceptance by humans, cost per item and time per item. When a step's pass rate is already near 100%, consider removing the gate; when a step fails often, split it further.

Key takeaways

  • Seven reusable patterns cover most LLM systems: chaining, routing, parallelization, orchestrator–workers, evaluator–optimizer, agent, multi-agent.
  • Prefer the simplest pattern that meets the quality bar; agents are the last resort, not the first.
  • Evaluator–optimizer loops need a specific rubric and a hard round cap.
  • Orchestrator–workers differs from parallelization because subtasks are discovered at runtime.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A support inbox receives refunds, technical questions and sales leads that need different handling. Which pattern fits first?
  2. What distinguishes orchestrator–workers from sectioning-style parallelization?
  3. Your evaluator–optimizer loop sometimes runs forever. What is the most important fix?

Put it into practice

Pick one process in your team and sketch it using at most three patterns from the catalog. Label where a code gate or human review sits.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.