Building AI Products & WorkflowsArchitecture, models and orchestration · Lesson 7 of 18

Workflow orchestration patterns for AI features

Article · 14 min · 9 min lecture

Video lecture

Workflow orchestration patterns for AI features

16 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 16

Workflow orchestration

  • Five recurring patterns
  • Choosing and combining
  • Production engineering

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Most AI products are workflows

Most production AI features are not autonomous agents. They are workflows: a sequence of steps, fixed in code or a workflow tool, in which models handle particular steps. Workflows are predictable, testable and cheap, and they cover the majority of business use cases. This lesson covers the five workflow patterns that recur everywhere (as popularised in Anthropic's widely cited guidance on building effective agents), and the engineering that makes them reliable in production.

Five patterns

1. Prompt chaining. Break a task into sequential steps, each a focused model call, with gates (checks in code) between them. Example: extract facts from a brief, then draft, then check the draft against brand rules.

2. Routing. Classify the input first, then send it to a specialised prompt, model or team. Example: support tickets routed to billing, technical or cancellation flows, with cheap models for easy routes.

3. Parallelisation. Run independent calls at the same time and combine results: sectioning (different parts of a task in parallel, such as reviewing a contract's clauses) or voting (the same task several times to increase confidence, such as flagging risky content).

4. Orchestrator-workers. A model decides the sub-tasks dynamically and delegates them, then synthesises. Useful when you cannot predict the sub-tasks in advance (this is where workflows start becoming agents).

5. Evaluator-optimiser. One call produces, another evaluates against explicit criteria and feeds back, for a bounded number of rounds. Example: translating marketing copy into Arabic, then checking tone and meaning against a rubric.

Choosing

SituationPattern
Steps are known and each is easier aloneChaining with gates
Inputs fall into distinct categoriesRouting
Independent sub-tasks, or you need confidenceParallelisation
Sub-tasks unknown until you see the inputOrchestrator-workers
Clear quality criteria, iteration helpsEvaluator-optimiser

Combine them freely: route first, then chain within each route, with an evaluator on the final step.

Production engineering

  • Gates in code between steps: schema validation, business rules, length limits, required citations. Fail fast and route to review rather than passing bad data forward.
  • Structured outputs between steps, so the next step receives validated data.
  • Idempotency and retries: each step can be retried safely; record step outputs so a retry resumes rather than restarts.
  • Durable execution for long or human-in-the-loop workflows: workflow engines (for example Temporal or cloud step-function services), queues with workers, or no-code platforms with wait and approval steps. A workflow waiting two days for a manager's approval must survive restarts.
  • Batch processing for non-urgent bulk steps, which many providers discount.
  • Observability: one trace per workflow run with a span per step, so you can see which step failed.

Hands-on: a chained workflow with gates and a batch variant

import os
from typing import Literal
import anthropic
from pydantic import BaseModel

client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")

class BriefFacts(BaseModel):
    brand: str
    audience: str
    key_messages: list[str]
    banned_claims: list[str]
    disclosure_required: bool

class Check(BaseModel):
    verdict: Literal["pass", "fail"]
    problems: list[str]

def step_extract(brief: str) -> BriefFacts:
    r = client.messages.parse(model=MODEL, max_tokens=1500, output_format=BriefFacts,
        messages=[{"role": "user", "content": f"Extract the campaign facts. Use only the brief.\n<brief>{brief}</brief>"}])
    return r.parsed_output

def gate_facts(f: BriefFacts):
    if not f.key_messages:
        raise ValueError("no key messages found: route to human")

def step_draft(f: BriefFacts) -> str:
    r = client.messages.create(model=MODEL, max_tokens=1200, messages=[{"role": "user", "content":
        f"Write 3 Instagram caption options for {f.brand} aimed at {f.audience}. Messages: {f.key_messages}. "
        f"Never use: {f.banned_claims}. {'Start with #ad.' if f.disclosure_required else ''}"}])
    return "".join(b.text for b in r.content if b.type == "text")

def step_check(f: BriefFacts, draft: str) -> Check:
    r = client.messages.parse(model=MODEL, max_tokens=800, output_format=Check, messages=[{"role": "user", "content":
        f"Check the captions. Fail if any banned claim appears {f.banned_claims} or, when required, #ad is missing "
        f"(required={f.disclosure_required}).\n<captions>{draft}</captions>"}])
    return r.parsed_output

def run(brief: str) -> dict:
    facts = step_extract(brief); gate_facts(facts)
    draft = step_draft(facts)
    for _ in range(2):                                  # evaluator-optimiser, bounded
        check = step_check(facts, draft)
        if check.verdict == "pass":
            return {"status": "ready_for_human_review", "draft": draft}
        draft = step_draft(facts) + f"\n\n(Previous problems to avoid: {check.problems})"
    return {"status": "needs_human", "draft": draft, "problems": check.problems}

For 500 briefs overnight, submit the extraction step through a batch endpoint instead, keyed by a custom_id per brief, and run later steps when results arrive:

batch = client.messages.batches.create(requests=[
    {"custom_id": f"brief-{i}", "params": {"model": MODEL, "max_tokens": 1500,
     "messages": [{"role": "user", "content": f"Extract campaign facts as JSON...\n<brief>{b}</brief>"}]}}
    for i, b in enumerate(briefs)])
# poll client.messages.batches.retrieve(batch.id).processing_status until "ended",
# then iterate client.messages.batches.results(batch.id) and match results by custom_id (order is not guaranteed)

A deterministic check (plain code looking for "#ad" and banned phrases) should back up the model-based check; use both.

Key takeaways

  • Most production AI features are workflows: fixed steps in code with models handling specific steps.
  • Five recurring patterns: prompt chaining, routing, parallelisation, orchestrator-workers and evaluator-optimiser.
  • Make workflows reliable with gates in code, structured outputs, idempotent retries, durable execution and per-step traces.
  • Use batch processing for non-urgent bulk steps and back model-based checks with deterministic ones.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Support tickets fall into billing, technical and cancellation categories that need different handling. Which pattern fits first?
  2. A workflow waits two days for a manager’s approval. What does it need?
  3. Why put gates in code between chained steps?

Put it into practice

Map one AI feature as a workflow: name each step, the pattern it uses, the gate after it, and where a human approves. Build the first two steps with structured outputs.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.