---
title: "Workflow and agent architectures: the pattern catalog"
description: "Why patterns matter Most successful LLM systems are built from a handful of composable patterns rather than sprawling frameworks. Knowing them lets you…"
url: https://optimizeall.com/learn/ai-agents-engineering/workflow-and-agent-patterns
updated: 2026-10-05
---

Building Production AI Agents · Agent foundations and architectures · lesson 2 of 18 · 16 min

# Workflow and agent architectures: the pattern catalog

## Why patterns matter

Most successful LLM systems are built from a handful of composable patterns rather than sprawling frameworks. Knowing them lets you pick the simplest design that works and explain it to stakeholders. The catalog below follows the widely used taxonomy popularized by Anthropic's "Building effective agents" guidance, extended with what teams have learned since.

## Pattern 1: Prompt chaining

Break a task into fixed sequential steps, each an LLM call, with optional **gates** (code checks) between them.

- Use when: the task decomposes cleanly (outline → draft → edit).
- Gate example: after generating an outline, check it has 4–7 sections before drafting.
- Trade-off: more latency, much higher reliability per step.

## Pattern 2: Routing

Classify the input, then send it to a specialized prompt, tool set or model.

- Use when: inputs fall into distinct categories with different handling (refund vs technical issue vs sales inquiry).
- Bonus: route easy queries to a small, cheap model and hard ones to a frontier model.
- Trade-off: misroutes; always measure classifier accuracy.

## Pattern 3: Parallelization

Run LLM calls concurrently and aggregate.

- **Sectioning**: split independent subtasks (analyze five competitor pages at once).
- **Voting**: run the same task several times and aggregate for confidence (for example, three independent reviews of whether an ad violates policy).
- Trade-off: cost scales with branches; latency drops.

## Pattern 4: Orchestrator–workers

A central model breaks the task into subtasks dynamically, delegates them to worker calls (often cheaper models), and synthesizes the results. Unlike parallelization, the subtasks are not known in advance.

- Use when: "Update our pricing mentions across the site" — you don't know which pages need changes until you look.

## Pattern 5: Evaluator–optimizer

One call generates, another evaluates against explicit criteria and returns feedback; loop until the evaluator passes or a limit is hit.

- Use when: clear evaluation criteria exist and iteration measurably improves output (translation nuance, ad copy meeting a brand checklist, code passing tests).
- Tip: the evaluator should be as specific as a rubric, not "is this good?".

## Pattern 6: The autonomous agent

A single model in a loop with tools, deciding its own steps until done. Simple to write, hard to make reliable. Invest in tool design, environment feedback (tests, validators) and stopping conditions.

## Pattern 7: Multi-agent systems

Several agents with separate contexts and roles: a lead agent with sub-agents, handoffs between specialists, or agents exposed to one another as tools. Powerful for breadth-first research and context isolation, but multi-agent systems typically consume many more tokens than a single agent, so reserve them for high-value tasks (module 2 goes deeper).

## Choosing: a decision table

| Situation | Start with |
|---|---|
| Known steps, quality matters | Prompt chaining with gates |
| Distinct input categories | Routing |
| Independent subtasks or need for confidence | Parallelization |
| Unknown subtasks discovered at runtime | Orchestrator–workers |
| Clear rubric and iteration helps | Evaluator–optimizer |
| Open-ended, many possible paths | Single agent with good tools |
| Very broad research or context overflow | Multi-agent (lead + sub-agents) |

## Worked example: a UK e-commerce brand's product-description pipeline

A London brand needs 400 product descriptions in brand voice, compliant with UK advertising rules (no unsupported claims like "clinically proven" without evidence).

1. **Routing**: classify each product (apparel, skincare, accessories). Skincare gets a stricter prompt about health claims.
2. **Prompt chaining**: extract attributes from the spec sheet → draft → gate (length 80–120 words, required keywords present).
3. **Evaluator–optimizer**: a checker prompt scores against a rubric (voice, claims, readability) and returns fixes; max two revision rounds.
4. **Parallelization**: process products concurrently within rate limits.

No agent needed. The result is predictable, cheap and auditable. The one place they did add an agent: a "research unusual ingredient" step that can search approved sources when the spec sheet mentions an unfamiliar component.

## Hands-on: an evaluator–optimizer in Python

```python
import os, json
import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from the environment
MODEL = os.environ.get("MODEL", "claude-sonnet-5")  # check current model IDs in the docs

RUBRIC = """Score 1-5 on each: brand_voice (warm, plain English), claims (no unsupported
health or performance claims), length (80-120 words). Return JSON:
{"scores": {...}, "pass": true|false, "fixes": ["..."]}. pass only if all scores >= 4."""

def ask(system: str, user: str) -> str:
    msg = client.messages.create(model=MODEL, max_tokens=2000, system=system,
                                 messages=[{"role": "user", "content": user}])
    return "".join(b.text for b in msg.content if b.type == "text")

def generate(spec: str, fixes: list[str] | None = None) -> str:
    extra = f"\nApply these fixes: {fixes}" if fixes else ""
    return ask("You write product descriptions for a UK skincare brand.",
               f"Spec sheet:\n{spec}{extra}\nWrite the description only.")

def evaluate(draft: str) -> dict:
    raw = ask("You are a strict copy reviewer. Output JSON only.", f"{RUBRIC}\n\nDraft:\n{draft}")
    return json.loads(raw[raw.find("{"): raw.rfind("}") + 1])

def pipeline(spec: str, max_rounds: int = 3) -> dict:
    draft, verdict = generate(spec), {}
    for round_no in range(max_rounds):
        verdict = evaluate(draft)
        if verdict.get("pass"):
            return {"draft": draft, "rounds": round_no + 1, "verdict": verdict}
        draft = generate(spec, verdict.get("fixes"))
    return {"draft": draft, "rounds": max_rounds, "verdict": verdict, "needs_human": True}
```

Notice the **round cap** and the **needs_human** flag: the loop always terminates, and failures route to a person. In production you would use structured outputs (covered later) instead of slicing JSON from text.

## Pitfalls

- Using an agent where routing plus chaining would do.
- Evaluators with vague criteria, which approve everything or loop forever.
- Frameworks that hide prompts; you must be able to see and version every prompt.
- Forgetting that parallel branches multiply rate-limit usage.

## Measuring success

For each pattern, track per-step pass rates (gates, evaluator verdicts), end-to-end acceptance by humans, cost per item and time per item. When a step's pass rate is already near 100%, consider removing the gate; when a step fails often, split it further.

## Video lecture: Workflow and agent architectures: the pattern catalog

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

1. Workflow and agent patterns
2. Why patterns matter
3. Chaining and routing
4. Parallel and orchestrator–workers
5. Evaluate, agent, multi-agent
6. Simple example: a newsletter helper
7. Example: 400 product descriptions
8. Hands-on: evaluator–optimizer
9. Pattern pitfalls
10. Decision guide
11. Deeper: 400 descriptions (illustrative)
12. Watch me do it: evaluator–optimizer
13. Try this now
14. Recap

## Lecture transcript

### Workflow and agent patterns

Here's a secret from teams shipping AI in production: most of the wins don't come from fancy autonomous agents. They come from a handful of simple, composable patterns. In this lesson you'll learn seven of them, when each one fits, and how to combine them so you always pick the simplest design that does the job.

### Why patterns matter

Why do patterns matter? Because the words agent and AI workflow get used for wildly different systems, and teams talk past each other. A shared pattern vocabulary fixes that. Think of it like cooking techniques. Every chef knows chop, sauté, braise and reduce. A recipe is just a combination of techniques. Once you know the seven patterns in this lesson, any AI system you meet, or any vendor demo you watch, becomes a recipe you can read: this part is routing, this part is chaining with a gate, this part is an evaluator loop. And you'll spot when someone is braising what should just be chopped.

### Chaining and routing

Pattern one is prompt chaining. You break a task into fixed steps, like outline, then draft, then edit, and you can put a code check, called a gate, between steps. For example, only draft if the outline has four to seven sections. Pattern two is routing. First classify the input, then send it to the right specialist prompt or model. A refund request and a technical question deserve different handling, and easy questions can go to a cheaper model.

### Parallel and orchestrator–workers

Pattern three is parallelization, which comes in two flavors. Sectioning splits independent work, like analyzing five competitor pages at the same time. Voting runs the same task several times and combines the answers, which is useful for judgment calls, like whether an ad breaks a policy. Pattern four is orchestrator workers. A lead model looks at the job, decides what subtasks exist, hands them to workers, and then combines the results. The key difference from sectioning: the subtasks aren't known until runtime.

### Evaluate, agent, multi-agent

Pattern five is evaluator optimizer. One call writes, another call judges against a clear rubric and sends back fixes. You loop until it passes or you hit a cap. This works brilliantly when criteria are explicit, like brand voice, no unsupported claims, and a word count. Pattern six is the single autonomous agent: a model in a loop with tools. Pattern seven is multi agent, where several agents with separate contexts collaborate. These last two are the most flexible, and also the most expensive and hardest to make reliable.

### Simple example: a newsletter helper

A simple example to lock this in. You run a small newsletter and want AI help. Step one, route: is today's input a product launch, an event or a tip? Each gets its own prompt. Step two, chain: write a headline, then a body, with a gate that checks the body is under two hundred words. Step three, evaluate: a reviewer prompt checks tone against three rules and sends back fixes, at most twice. Three patterns, zero agents, and a result you can predict. Notice how each pattern solves one specific problem. That's the mindset.

### Example: 400 product descriptions

Now a worked example. A London skincare brand needs four hundred product descriptions in brand voice that also respect UK advertising rules, so no unsupported claims like clinically proven. They route each product by category, with skincare getting a stricter prompt. They chain extract, draft and a length gate. They add an evaluator with a rubric and a two round cap. And they process items in parallel within rate limits. Notice there is no agent at all, and the result is predictable, cheap and auditable.

### Hands-on: evaluator–optimizer

In the lesson text you'll find a short Python evaluator optimizer. Look for two details. First, the loop has a maximum number of rounds, so it always ends. Second, if it never passes, the result is flagged for a human. Those two lines matter more than any clever prompt. Also notice that every prompt is visible in your code. If a framework hides prompts from you, you can't version, test or debug them.

### Pattern pitfalls

Let's look at the traps. The first is reaching for an agent when routing plus a couple of chained steps would do. The second is an evaluator with a vague rubric. If you ask, is this good, it will either approve everything or send you round in circles. The third is a framework that hides the prompts, so you can't see what's actually being sent. And the fourth is forgetting that parallel branches multiply your rate limit usage. Ten branches means ten times the requests per minute, so throttle your fan out to stay inside your limits.

### Decision guide

A quick decision guide. Known steps and quality matters: chaining. Different input types: routing. Independent subtasks: parallel. Unknown subtasks: orchestrator workers. Clear rubric: evaluator optimizer. Open ended with many paths: an agent. Research that's too broad for one context window: multi agent. And measure every step. If a gate passes almost everything, maybe remove it. If a step fails often, split it further.

### Deeper: 400 descriptions (illustrative)

Let's deepen the London skincare example with illustrative numbers. Four hundred descriptions, three categories. Routing sends about one hundred and twenty skincare items to the stricter prompt. The length gate catches around one in ten first drafts, which are regenerated automatically. The evaluator passes most drafts on the first round, fixes the rest on the second, and flags about a dozen for a human, mostly products whose spec sheets used medical sounding language. A copywriter reviews those twelve instead of four hundred. The whole batch runs overnight at a small fraction of a freelancer's fee, and, crucially, every step has a log the brand's compliance lead can read. That audit trail is what made legal comfortable enough to approve the project.

### Watch me do it: evaluator–optimizer

Watch me do it. I'll walk through the evaluator optimizer code line by line. At the top, the rubric: three scores from one to five, brand voice, claims and length, returned as JSON with a pass flag and a list of fixes. The ask function wraps one Claude call and joins the text blocks. Generate takes the spec sheet and, optionally, the fixes from the last review. Evaluate sends the draft and the rubric to a strict reviewer prompt and pulls the JSON out of the reply. Now the pipeline. It generates a first draft, then loops at most three times. Each round, it evaluates. If the verdict passes, it returns the draft and the number of rounds. If not, it regenerates with the fixes. And if it runs out of rounds, it returns the last draft with needs human set to true. Two details matter most: the loop can never run forever, and failures land with a person instead of silently shipping.

### Try this now

Try this now. Pick one real process from your week, like onboarding a client, producing a report or answering a type of support ticket. Draw it as boxes and arrows on paper. Then label each box with one of the seven patterns. Where would a code gate catch errors? Where does a human review? Where, if anywhere, does the path genuinely depend on what the model discovers? If you can describe the whole thing with chaining, routing and one evaluator, you've just designed a reliable system without an agent. Keep this sketch; you'll use it again in the capstone.

### Recap

To recap: seven patterns cover almost every LLM system. Start with the simplest one that meets your quality bar, cap every loop, keep prompts visible, and measure each step. Your next step: take one real process from your team and sketch it with no more than three patterns, marking where a code gate or a human review sits. You'll reuse that sketch in the capstone.

## Key takeaways

- Seven reusable patterns cover most LLM systems: chaining, routing, parallelization, orchestrator–workers, evaluator–optimizer, agent, multi-agent.
- Prefer the simplest pattern that meets the quality bar; agents are the last resort, not the first.
- Evaluator–optimizer loops need a specific rubric and a hard round cap.
- Orchestrator–workers differs from parallelization because subtasks are discovered at runtime.

## Try it

Pick one process in your team and sketch it using at most three patterns from the catalog. Label where a code gate or human review sits.

- [Previous: What an AI agent really is (and when not to build one)](https://optimizeall.com/learn/ai-agents-engineering/what-is-an-agent)
- [Next: Build the agent loop from scratch in Python](https://optimizeall.com/learn/ai-agents-engineering/agent-loop-from-scratch)
- [All lessons of Building Production AI Agents](https://optimizeall.com/learn/ai-agents-engineering)
