---
title: "Decomposition and prompt chaining | Optimize All Academy"
description: "One prompt, many jobs: the hidden problem When a single prompt asks the model to research, analyse, decide, write and format, quality usually drops on at…"
url: https://optimizeall.com/learn/advanced-prompt-engineering/decomposition-and-prompt-chaining
updated: 2026-10-05
---

Advanced Prompt Engineering · Decomposition, reasoning and verification · lesson 5 of 17 · 12 min

# Decomposition and prompt chaining

## One prompt, many jobs: the hidden problem

When a single prompt asks the model to research, analyse, decide, write and format, quality usually drops on at least one of those jobs, and you cannot tell which. Decomposition splits a complex task into subtasks, each with its own focused prompt and its own check. Prompt chaining passes the output of one step as input to the next.

The mental model is an assembly line with quality control stations: every station does one thing well and hands over a well-defined artefact.

## When to decompose

Decompose when:

- The task has **distinct phases** with different skills (extraction vs persuasion writing).
- You need to **inspect or control intermediate results** (for example approve an outline before drafting).
- One phase needs **different context** (the drafting step does not need the 50 raw reviews, only the extracted themes).
- You want to **use different models** per step (a fast, cheap model for extraction, a stronger one for synthesis).

Do not decompose when the task is simple enough that one clear prompt performs well; every extra step adds latency, cost and a place for errors to compound.

## Common chain patterns

1. **Sequential pipeline.** Extract, then analyse, then write, then format.
2. **Map-reduce.** Apply the same prompt to many chunks in parallel, then combine.
3. **Router.** A first classification step sends input to one of several specialised prompts (billing, technical, sales).
4. **Generate-then-refine.** Draft, critique against criteria, revise (covered in the verification lesson).
5. **Branch-and-select.** Generate several candidates, then score and choose.

## A worked example: turning customer reviews into a product brief

A single prompt ("Read these 300 reviews and write a product improvement brief") yields a generic brief. A chained version:

```text
Step 1  EXTRACT   (per batch of 25 reviews, in parallel)
        Output JSON: [{theme, sentiment, quote, review_id}]
Step 2  CLUSTER   (all extracted items)
        Merge duplicates, count mentions per theme, keep 2 best quotes each
Step 3  PRIORITISE
        Rank themes by frequency x severity; flag anything safety-related
Step 4  WRITE
        Produce a 1-page brief for the product team from the ranked themes
Step 5  CHECK (code + model)
        Code: every quote in the brief exists in the source reviews
        Model: does the brief overstate any finding relative to counts?
```

Each step has a clear input and output. If the brief is wrong, you can see whether extraction missed things, clustering merged unrelated themes, or writing exaggerated.

## Designing the hand-offs

The quality of a chain lives in its interfaces:

- **Structured intermediates.** Pass JSON or tagged sections, not loose prose, between steps.
- **Carry provenance.** Keep IDs (review_id, document page) through every step so final claims can be traced.
- **Minimal context per step.** Give each step only what it needs; the writing step does not need raw reviews.
- **Deterministic glue in code.** Counting, sorting, deduplicating exact matches and arithmetic should be code, not model calls. Use models for judgement, language and fuzzy matching.

## Error compounding

If each of five steps is 95% reliable and errors are independent, the whole chain is only about 77% reliable (0.95 multiplied by itself five times). This illustrative calculation is why chains need checks at the riskiest steps, and why fewer, well-designed steps can beat many small ones.

## Chains vs agents

A chain has a fixed sequence that you design. An agent decides its own next step in a loop. Chains are more predictable, cheaper and easier to test; prefer them whenever the steps are known in advance. Reach for agent-style loops only when the path genuinely depends on what is discovered along the way. Our course on RAG, tool use and agents explores that boundary in depth.

## Failure modes

- **Telephone game.** Meaning degrades as each step paraphrases the previous one. Pass original quotes, not summaries of summaries.
- **Silent step failure.** Step 2 returns an empty list, and step 4 cheerfully writes a brief anyway. Add guards: if an intermediate is empty or malformed, stop and alert.
- **Over-engineering.** Ten-step chains for a task a single prompt handles well. Always compare against a single-prompt baseline.

## Hands-on: a three-step chain with guards

This chain extracts themes from review batches in parallel, merges them in code, then writes a brief. Each step has a guard that stops the chain if the intermediate is empty or malformed.

```python
import json, os, concurrent.futures as cf
import anthropic

client = anthropic.Anthropic()
MODEL = os.environ.get("CLAUDE_MODEL", "claude-opus-5")
FAST_MODEL = os.environ.get("CLAUDE_FAST_MODEL", "claude-haiku-4-5")  # cheaper step

THEMES_SCHEMA = {
    "type": "object",
    "properties": {"items": {"type": "array", "items": {
        "type": "object",
        "properties": {"review_id": {"type": "string"}, "theme": {"type": "string"},
                       "sentiment": {"type": "string", "enum": ["pos", "neg", "mixed"]},
                       "quote": {"type": "string"}},
        "required": ["review_id", "theme", "sentiment", "quote"],
        "additionalProperties": False}}},
    "required": ["items"], "additionalProperties": False,
}

def extract(batch):
    text = "\n".join(f'<review id="{r["id"]}">{r["text"]}</review>' for r in batch)
    resp = client.messages.create(
        model=FAST_MODEL, max_tokens=4096,
        output_config={"format": {"type": "json_schema", "schema": THEMES_SCHEMA}},
        messages=[{"role": "user", "content": "Extract one item per theme mentioned. "
                   "Quote exactly.\n" + text}],
    )
    return json.loads(next(b.text for b in resp.content if b.type == "text"))["items"]

def run(reviews, batch_size=25):
    batches = [reviews[i:i + batch_size] for i in range(0, len(reviews), batch_size)]
    with cf.ThreadPoolExecutor(max_workers=4) as pool:
        items = [it for part in pool.map(extract, batches) for it in part]
    if not items:
        raise RuntimeError("Extraction returned nothing; stopping before the write step")
    source = {r["id"]: r["text"] for r in reviews}
    items = [it for it in items if it["quote"] in source.get(it["review_id"], "")]  # provenance guard
    counts = {}
    for it in items:
        counts.setdefault(it["theme"].lower(), []).append(it)
    ranked = sorted(counts.items(), key=lambda kv: -len(kv[1]))[:8]
    brief_input = json.dumps([{"theme": t, "mentions": len(v), "quotes": [x["quote"] for x in v[:2]]}
                              for t, v in ranked], ensure_ascii=False)
    resp = client.messages.create(
        model=MODEL, max_tokens=2048,
        messages=[{"role": "user", "content": "Write a one-page product brief from these ranked "
                   "themes. Do not overstate: use the mention counts as given.\n" + brief_input}],
    )
    return "".join(b.text for b in resp.content if b.type == "text")
```

Notice the division of labour: models extract and write; code merges, counts, ranks and verifies quotes. Exact-match theme merging is deliberately naive here; for production, add a clustering step and log the merges.

## Use the right model per step

Chains let you right-size models: a small, fast model for high-volume extraction or routing and a stronger model for synthesis. Measure each step separately on your eval set; if the cheap model's extraction misses themes, the expensive writer cannot recover them.

## Measuring a chain

Log each step's input, output, tokens and latency with a run ID. Track per-step failure rates, not only the final score. When the final output is wrong, you should be able to point to the step that caused it within minutes.

## Going further

Log every intermediate artefact with a run ID. When a stakeholder questions an output weeks later, you can replay the chain and see exactly where a claim came from. This is also your best dataset for improving individual steps.

## Video lecture: Decomposition and prompt chaining

Lecture coming soon · 11 chapters · about 8 minutes. Read the full transcript below.

1. Decomposition and prompt chaining
2. When to decompose
3. Five chain patterns
4. Worked example: reviews to brief
5. Designing hand-offs
6. Errors compound
7. Right-size and log
8. Chains vs agents
9. Example 1: podcast → blog post
10. Example 2: survey analysis (illustrative)
11. Recap

## Lecture transcript

### Decomposition and prompt chaining

Read these three hundred reviews and write a product improvement brief. It sounds like one task. It is actually five: extract, cluster, prioritise, write and check. When one prompt does all five, quality drops on at least one of them, and you cannot tell which. In this lecture you will learn when to decompose a task, the five chain patterns worth knowing, how to design the hand-offs between steps, why errors compound, and when a chain beats an agent.

### When to decompose

Think of a chain as an assembly line with quality control stations. Every station does one thing well and hands over a well-defined artefact. Decompose when the task has distinct phases that need different skills, like extraction versus persuasive writing. When you need to inspect or approve intermediate results, like an outline before drafting. When one phase needs different context, since the writing step does not need fifty raw reviews. Or when you want different models per step, a fast, cheap model for extraction and a stronger one for synthesis. Do not decompose simple tasks; every step adds latency, cost and another place for errors.

### Five chain patterns

Five chain patterns cover most needs. A sequential pipeline: extract, then analyse, then write, then format. Map-reduce: the same prompt applied to many chunks in parallel, then combined. A router: a first classification step sends each input to a specialised prompt, like billing, technical or sales. Generate-then-refine: draft, critique against criteria, revise. And branch-and-select: generate several candidates, then score and choose the best.

### Worked example: reviews to brief

Here is the review brief as a chain. Step one extracts themes from batches of twenty-five reviews in parallel, returning JSON with theme, sentiment, an exact quote and the review ID. Step two clusters: merge duplicates, count mentions per theme, keep the two best quotes. Step three prioritises by frequency and severity, flagging anything safety-related. Step four writes a one-page brief from the ranked themes. Step five checks: code confirms every quote in the brief exists in the source reviews, and a model check asks whether the brief overstates any finding relative to the counts. If the brief is wrong, you can see exactly which step caused it.

### Designing hand-offs

The quality of a chain lives in its interfaces. Pass structured intermediates, JSON or tagged sections, not loose prose. Carry provenance: keep review IDs or page numbers through every step, so final claims can be traced. Give each step only the context it needs. And put deterministic glue in code: counting, sorting, exact deduplication and arithmetic should be code, not model calls. Use models for judgement, language and fuzzy matching. In the lesson's hands-on code, models extract and write, while plain Python merges, counts, ranks, and drops any quote that does not appear verbatim in its source review.

### Errors compound

Now the maths that should make you careful. If each of five steps is ninety-five percent reliable, and errors are independent, the whole chain is only about seventy-seven percent reliable. That is ninety-five percent multiplied by itself five times. It is an illustrative calculation, but the lesson is real: chains need checks at the riskiest steps, and fewer, well-designed steps can beat many small ones. Also add guards: if an intermediate result is empty or malformed, stop and alert, instead of cheerfully writing a brief from nothing.

### Right-size and log

Chains also let you right-size models. A small, fast model can handle high-volume extraction or routing, while a stronger model handles the synthesis. But measure each step separately on your eval set. If the cheap model's extraction misses themes, the expensive writer cannot recover them. And log every step's input, output, tokens and latency with a run ID. When a stakeholder questions an output weeks later, you can replay the chain and see exactly where a claim came from.

### Chains vs agents

What about agents? A chain has a fixed sequence that you design. An agent decides its own next step in a loop. Chains are more predictable, cheaper and easier to test, so prefer them whenever the steps are known in advance. Reach for agent-style loops only when the path genuinely depends on what is discovered along the way. And watch three chain failure modes: the telephone game, where meaning degrades as each step paraphrases the last, so pass original quotes; silent step failures; and over-engineering, ten steps for a task one prompt handles well. Always compare against a single-prompt baseline.

### Example 1: podcast → blog post

A simple worked example. You want a blog post from a podcast transcript. One prompt, write a blog post from this transcript, gives a vague summary. A three-step chain works better. Step one: extract the five most interesting ideas, each with a direct quote and a timestamp. Step two, which is you: pick the three ideas you want. Step three: write an eight-hundred-word post built around those three ideas, using the quotes. Each step is simple, you get a checkpoint in the middle, and every quote in the final post can be traced to a timestamp. Notice that the human choice in the middle is part of the chain. Chains do not have to be fully automated.

### Example 2: survey analysis (illustrative)

Now a business scenario, with illustrative numbers. A market research agency in London analyses about fifteen hundred open-ended survey responses per client project. Their single-prompt approach produced reports with inflated claims, like most customers want this, based on a handful of comments. They built a four-step chain. A small, fast model extracts themes from batches of fifty responses, with response IDs and exact quotes. Code merges and counts themes. A stronger model writes the report from the counts. And a check step verifies every quote exists in the source and that no claim exceeds the counted share. Extraction runs in parallel, so the whole chain finishes in minutes. Cost per project fell by around half because most tokens now go through the small model, and client challenges about overstated findings largely disappeared. Illustrative figures, but the structure is the key idea.

### Recap

To recap. Decompose when phases differ, when you need inspection points, different context or different models. Use the five patterns, design structured hand-offs with provenance, and put counting and arithmetic in code. Expect errors to compound, add guards and checks, and prefer chains to agents when the steps are known. Try this now: map one multi-step task you currently run as a single prompt into a three to five step chain, and write the input, output format and one automated check for each step. Next: reasoning models, and when step-by-step thinking helps or hurts.

## Key takeaways

- Decompose when a task has distinct phases, needs inspection between steps, different context or different models.
- Common patterns: sequential pipeline, map-reduce, router, generate-then-refine and branch-and-select.
- Chain quality lives in hand-offs: structured intermediates, provenance IDs and deterministic glue in code.
- Errors compound across steps, so add checks at risky steps and compare against a single-prompt baseline.

## Try it

Map one multi-step task you currently do with a single prompt into a 3-5 step chain. For each step write its input, output format and one automated check.

- [Previous: Structured outputs and JSON schemas](https://optimizeall.com/learn/advanced-prompt-engineering/structured-outputs-json-schemas)
- [Next: Reasoning models: when step-by-step helps and when it hurts](https://optimizeall.com/learn/advanced-prompt-engineering/reasoning-models-step-by-step)
- [All lessons of Advanced Prompt Engineering](https://optimizeall.com/learn/advanced-prompt-engineering)
