Building AI Products & WorkflowsArchitecture, models and orchestration · Lesson 7 of 18
Workflow orchestration patterns for AI features
Video lecture
Workflow orchestration patterns for AI features
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Workflow orchestration
Agents get the headlines, but most AI features that quietly save businesses hours every day are workflows. Fixed steps, defined in code or a workflow tool, with a model doing specific jobs inside them. They're predictable, testable and cheap. In this lesson you'll learn the five workflow patterns that show up everywhere, how to choose between them, and the production engineering, gates, retries, durable execution and batch processing, that makes them reliable.
0:31 Analogy: the kitchen line
Here's an analogy. A workflow is like a restaurant kitchen's line. Orders come in, a station preps, another cooks, another plates, and the pass checks each plate before it leaves. Routing is the head chef deciding which station gets the order. Parallelisation is several cooks working at once. The evaluator-optimiser is the pass, sending a plate back to be fixed. The model is a talented cook at some stations, not the whole kitchen.
1:03 Why it matters
Why does this matter for your product? Because workflows are where most of the reliability, cost savings and speed come from. Agents grab attention, but a well-designed workflow with two focused model calls and good checks will usually beat a clever agent on cost, predictability and ease of testing. Learn these patterns and you'll ship more features that actually stay working.
1:30 Patterns 1–2
Pattern one: prompt chaining. Break a task into steps, each a focused model call, with gates between them. Extract facts from a brief, then draft, then check the draft against brand rules. Each step is easier alone, and gates stop bad data flowing forward. Pattern two: routing. Classify the input first, then send it down a specialised path. Billing, technical and cancellation tickets each get their own prompt, model or team, and easy routes can use cheaper models.
2:04 Patterns 3–5
Pattern three: parallelisation. Run independent calls at the same time. Sectioning splits a task, like reviewing each clause of a contract in parallel. Voting runs the same task several times to boost confidence, like flagging risky content only when most runs agree. Pattern four: orchestrator-workers. A model decides the sub-tasks based on the input, delegates them and combines results. It's useful when you can't predict the sub-tasks, and it's where workflows start to become agents. Pattern five: evaluator-optimiser. One call produces, another evaluates against clear criteria and feeds back, for a limited number of rounds.
2:45 Choosing and combining
Choose by situation. Known steps that are easier separately: chaining. Distinct categories of input: routing. Independent sub-tasks, or a need for confidence: parallelisation. Sub-tasks you can't know in advance: orchestrator-workers. Clear quality criteria where iteration helps, like translating marketing copy into Arabic and checking tone and meaning against a rubric: evaluator-optimiser. And combine them. Route first, chain within each route, and put an evaluator on the final step.
3:15 Simple example: two steps, two gates
A simple example of chaining with a gate. Step one: extract the order number and complaint type from a customer email. Gate: is the order number in the right format and does it exist? If not, send it to a person. Step two: draft a reply using the order's details. Gate: does the draft mention a refund amount? If so, route it for approval. Two steps, two gates, and each problem is caught exactly where it starts.
3:48 Production engineering
Now the engineering. Put gates in code between steps: schema validation, business rules, length limits and required citations, and route failures to review. Pass structured outputs between steps. Make each step idempotent and retryable, and save step outputs so a retry resumes instead of restarting. For long or human-in-the-loop workflows, use durable execution: a workflow engine, a queue with workers, or a no-code platform with wait and approval steps. A workflow waiting two days for a manager must survive a server restart.
4:24 Batch + observability
Two more practices. Batch processing: for non-urgent bulk work, like extracting facts from five hundred briefs overnight, batch endpoints are often discounted, and results come back keyed by an ID you choose, not in order. And observability: one trace per workflow run, with a span per step, so when something fails you can see exactly which step and why. Also, back up model-based checks with deterministic ones. If a caption must contain hashtag ad, a line of code can check that with certainty.
5:00 Business example (illustrative)
A deeper business example, illustrative. A London creator-management agency runs this caption workflow for about three hundred sponsored posts a month. The check step catches a missing disclosure or banned claim in roughly one draft in twelve, and the deterministic check catches the few the model misses. Account managers now review drafts that are almost always compliant, and time per sponsored post fell from around forty minutes to fifteen.
5:30 Hands-on in the lesson
The hands-on section builds a creator-marketing workflow: extract campaign facts from a brief with a structured output, gate them in code, draft Instagram captions that avoid banned claims and include disclosure when required, then check them with a second structured call, retrying at most twice before handing to a human. You'll also see how to send the extraction step through a batch endpoint for five hundred briefs overnight, matching results by custom ID.
6:02 Common mistakes
Common mistakes. Using one giant prompt to do extraction, drafting and checking all at once. Putting no checks between steps. Retrying failed steps from the very beginning, so earlier work is repeated and sometimes duplicated. Running long waits inside a single process that dies on restart. And evaluating only the final output, so you can't see which step caused a failure.
6:29 How you'll know it's working
How will you know your workflow is well orchestrated? Each step has its own pass rate you can see. Failures show up at the gate where they start, not at the end. A crash or restart doesn't lose or duplicate work. Human waits don't block anything else. And cost per completed run is stable as volume grows, with batch processing used where speed isn't needed.
6:57 Watch me do it: chain with gates
Watch me do it. I open the workflow code. First, two schemas: BriefFacts with brand, audience, key messages, banned claims and whether disclosure is required, and Check with a pass or fail verdict and a list of problems. Step extract parses the brief into BriefFacts. The gate raises an error if there are no key messages, which routes the brief to a human. Step draft writes three captions using the messages, avoiding banned claims and starting with the disclosure when required. Step check parses a verdict. Then run loops at most twice: if the check passes, it returns ready for human review; if not, it redrafts with the problems as guidance. After two failures it returns needs human. I run it on a skincare brief with a banned claim, clinically proven. The first draft fails, the second passes, and the status is ready for human review.
8:00 Recap
To recap: most AI features are workflows. Use chaining, routing, parallelisation, orchestrator-workers and evaluator-optimiser deliberately, and combine them. Make them production-grade with gates, structured outputs, idempotent retries, durable execution, batching and traces. Your next step is to map one AI feature as a workflow, naming each step's pattern, gate and human checkpoint, and build the first two steps. Next, integrating AI into the systems you already run.
8:29 Try this now (45 minutes)
Try this now. Take one AI feature and draw it as boxes: each step, its pattern, the gate after it, and where a person approves. Mark any step that could run in parallel and any that could run overnight in a batch. Then build the first two steps with structured outputs and one gate in code, and run ten real inputs through them.
Most AI products are workflows
Most production AI features are not autonomous agents. They are workflows: a sequence of steps, fixed in code or a workflow tool, in which models handle particular steps. Workflows are predictable, testable and cheap, and they cover the majority of business use cases. This lesson covers the five workflow patterns that recur everywhere (as popularised in Anthropic's widely cited guidance on building effective agents), and the engineering that makes them reliable in production.
Five patterns
1. Prompt chaining. Break a task into sequential steps, each a focused model call, with gates (checks in code) between them. Example: extract facts from a brief, then draft, then check the draft against brand rules.
2. Routing. Classify the input first, then send it to a specialised prompt, model or team. Example: support tickets routed to billing, technical or cancellation flows, with cheap models for easy routes.
3. Parallelisation. Run independent calls at the same time and combine results: sectioning (different parts of a task in parallel, such as reviewing a contract's clauses) or voting (the same task several times to increase confidence, such as flagging risky content).
4. Orchestrator-workers. A model decides the sub-tasks dynamically and delegates them, then synthesises. Useful when you cannot predict the sub-tasks in advance (this is where workflows start becoming agents).
5. Evaluator-optimiser. One call produces, another evaluates against explicit criteria and feeds back, for a bounded number of rounds. Example: translating marketing copy into Arabic, then checking tone and meaning against a rubric.
Choosing
| Situation | Pattern |
|---|---|
| Steps are known and each is easier alone | Chaining with gates |
| Inputs fall into distinct categories | Routing |
| Independent sub-tasks, or you need confidence | Parallelisation |
| Sub-tasks unknown until you see the input | Orchestrator-workers |
| Clear quality criteria, iteration helps | Evaluator-optimiser |
Combine them freely: route first, then chain within each route, with an evaluator on the final step.
Production engineering
- Gates in code between steps: schema validation, business rules, length limits, required citations. Fail fast and route to review rather than passing bad data forward.
- Structured outputs between steps, so the next step receives validated data.
- Idempotency and retries: each step can be retried safely; record step outputs so a retry resumes rather than restarts.
- Durable execution for long or human-in-the-loop workflows: workflow engines (for example Temporal or cloud step-function services), queues with workers, or no-code platforms with wait and approval steps. A workflow waiting two days for a manager's approval must survive restarts.
- Batch processing for non-urgent bulk steps, which many providers discount.
- Observability: one trace per workflow run with a span per step, so you can see which step failed.
Hands-on: a chained workflow with gates and a batch variant
import os
from typing import Literal
import anthropic
from pydantic import BaseModel
client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
class BriefFacts(BaseModel):
brand: str
audience: str
key_messages: list[str]
banned_claims: list[str]
disclosure_required: bool
class Check(BaseModel):
verdict: Literal["pass", "fail"]
problems: list[str]
def step_extract(brief: str) -> BriefFacts:
r = client.messages.parse(model=MODEL, max_tokens=1500, output_format=BriefFacts,
messages=[{"role": "user", "content": f"Extract the campaign facts. Use only the brief.\n<brief>{brief}</brief>"}])
return r.parsed_output
def gate_facts(f: BriefFacts):
if not f.key_messages:
raise ValueError("no key messages found: route to human")
def step_draft(f: BriefFacts) -> str:
r = client.messages.create(model=MODEL, max_tokens=1200, messages=[{"role": "user", "content":
f"Write 3 Instagram caption options for {f.brand} aimed at {f.audience}. Messages: {f.key_messages}. "
f"Never use: {f.banned_claims}. {'Start with #ad.' if f.disclosure_required else ''}"}])
return "".join(b.text for b in r.content if b.type == "text")
def step_check(f: BriefFacts, draft: str) -> Check:
r = client.messages.parse(model=MODEL, max_tokens=800, output_format=Check, messages=[{"role": "user", "content":
f"Check the captions. Fail if any banned claim appears {f.banned_claims} or, when required, #ad is missing "
f"(required={f.disclosure_required}).\n<captions>{draft}</captions>"}])
return r.parsed_output
def run(brief: str) -> dict:
facts = step_extract(brief); gate_facts(facts)
draft = step_draft(facts)
for _ in range(2): # evaluator-optimiser, bounded
check = step_check(facts, draft)
if check.verdict == "pass":
return {"status": "ready_for_human_review", "draft": draft}
draft = step_draft(facts) + f"\n\n(Previous problems to avoid: {check.problems})"
return {"status": "needs_human", "draft": draft, "problems": check.problems}For 500 briefs overnight, submit the extraction step through a batch endpoint instead, keyed by a custom_id per brief, and run later steps when results arrive:
batch = client.messages.batches.create(requests=[
{"custom_id": f"brief-{i}", "params": {"model": MODEL, "max_tokens": 1500,
"messages": [{"role": "user", "content": f"Extract campaign facts as JSON...\n<brief>{b}</brief>"}]}}
for i, b in enumerate(briefs)])
# poll client.messages.batches.retrieve(batch.id).processing_status until "ended",
# then iterate client.messages.batches.results(batch.id) and match results by custom_id (order is not guaranteed)A deterministic check (plain code looking for "#ad" and banned phrases) should back up the model-based check; use both.
Key takeaways
- Most production AI features are workflows: fixed steps in code with models handling specific steps.
- Five recurring patterns: prompt chaining, routing, parallelisation, orchestrator-workers and evaluator-optimiser.
- Make workflows reliable with gates in code, structured outputs, idempotent retries, durable execution and per-step traces.
- Use batch processing for non-urgent bulk steps and back model-based checks with deterministic ones.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Map one AI feature as a workflow: name each step, the pattern it uses, the gate after it, and where a human approves. Build the first two steps with structured outputs.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.