Building Production AI AgentsMemory, retrieval and durable state · Lesson 9 of 18

State, durability and long-running agents

Article · 15 min · 9 min lecture

Video lecture

State, durability and long-running agents

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

State and durability

  • What counts as agent state
  • Checkpoints and idempotency
  • Execution models for long runs

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why durability becomes the problem

A demo agent runs in one process for 30 seconds. A production agent may run for minutes or hours, wait for human approval overnight, survive deploys and crash mid-step. Without durable state, a crash means lost work, duplicated side effects (two refunds, two emails), or a confused user. Durability is the property that an agent's progress survives failures and can resume exactly where it stopped.

What counts as agent state

  • The message history (including tool calls and results).
  • The plan or todo list and which steps are complete.
  • Structured facts (budgets, IDs, approvals granted).
  • Pending actions awaiting approval.
  • Budgets consumed (steps, tokens, money).
  • References to external artifacts (files, drafts, tickets).

Store it outside the process: a database row per run, keyed by run_id, updated after every step.

Checkpointing

After every model call and tool execution, persist a checkpoint. On restart, load the latest checkpoint and continue the loop. Frameworks do this for you: LangGraph compiles graphs with a checkpointer (in-memory for dev; Postgres or others for production) and supports interrupt() for pauses; the OpenAI Agents SDK has sessions and serialisable run state for resuming after approvals; managed agent platforms keep sessions server-side. Durable-execution engines such as Temporal are also used to wrap agent steps as retriable activities.

Idempotency: the rule that prevents double actions

Any tool with side effects must be safe to retry. Pattern:

  1. Generate an idempotency key for each intended action (for example run_id + step + tool_name).
  2. Pass it to the downstream API (many payment and messaging APIs accept one) or record it in your own table before executing.
  3. If the key already exists with a result, return the stored result instead of acting again.

This matters because retries happen at every layer: SDK retries, queue redelivery, your own resume logic.

Execution models for long runs

ModelHow it worksUse when
Synchronous requestAPI call waits for the agent to finishRuns under ~30–60 s, interactive chat
StreamingStream progress events to the UIInteractive but multi-step
Background jobEnqueue run; worker processes; client polls or gets webhookMinutes-long research, batch work
ScheduledCron triggers runsDaily reports, monitoring
Event-drivenWebhooks (new lead, new ticket) start runsOps automation

Background jobs need a queue (e.g., a managed queue service or Redis-backed workers), a runs table, and status endpoints.

Worked example: overnight competitor monitoring for a US DTC brand

Every night at 02:00 local time, an agent checks 30 competitor sites for price changes, drafts a summary and proposes price responses for approval.

  • Scheduled trigger creates run_id, enqueues one job per competitor (fan-out) plus a final synthesis job.
  • Each job checkpoints after every tool call; a crash at competitor 17 resumes at 17.
  • Proposed price changes are stored as pending actions with idempotency keys; the pricing manager approves in the morning; execution happens once even if the approval webhook fires twice.
  • Budgets: 12 tool calls per competitor, a nightly token cap, alert if exceeded.

Hands-on: a resumable loop with a runs table

import json, sqlite3, uuid

db = sqlite3.connect("runs.db")
db.execute("""CREATE TABLE IF NOT EXISTS runs(
  run_id TEXT PRIMARY KEY, status TEXT, step INTEGER, messages TEXT, updated REAL)""")
db.execute("""CREATE TABLE IF NOT EXISTS actions(
  idem_key TEXT PRIMARY KEY, result TEXT)""")

def save(run_id, status, step, messages):
    db.execute("REPLACE INTO runs VALUES (?,?,?,?,strftime('%s','now'))",
               (run_id, status, step, json.dumps(messages, default=str)))
    db.commit()

def load(run_id):
    row = db.execute("SELECT status, step, messages FROM runs WHERE run_id=?", (run_id,)).fetchone()
    return (row[0], row[1], json.loads(row[2])) if row else None

def once(idem_key: str, action):
    """Run a side-effecting action at most once per key."""
    row = db.execute("SELECT result FROM actions WHERE idem_key=?", (idem_key,)).fetchone()
    if row:
        return json.loads(row[0])
    result = action()
    db.execute("INSERT INTO actions VALUES (?,?)", (idem_key, json.dumps(result)))
    db.commit()
    return result

def start(goal: str) -> str:
    run_id = str(uuid.uuid4())
    save(run_id, "running", 0, [{"role": "user", "content": goal}])
    return run_id

# In the loop from lesson 3: after each step call save(run_id, "running", step, messages).
# Before side effects: once(f"{run_id}:{step}:{tool_name}", lambda: do_the_thing(**args)).
# On worker start: status, step, messages = load(run_id) and continue from there.

Note that SDK response objects must be serialized to plain dicts (for example with .model_dump() on Anthropic/OpenAI Pydantic models) before storing.

Pitfalls

  • Holding state only in memory in a web process that autoscaling can kill.
  • Retrying side-effecting tools without idempotency keys.
  • Very long synchronous HTTP requests that time out at load balancers.
  • No status endpoint, so users refresh and start duplicate runs.

Measuring success

Track resumption success after injected failures (kill a worker mid-run in staging), duplicate-action rate (target zero), run completion rate, and time-to-resume. Chaos-test durability before you trust it.

Key takeaways

  • Durability means an agent's progress survives crashes, deploys and overnight waits.
  • Persist a checkpoint after every model call and tool execution, keyed by run ID.
  • Every side-effecting tool needs idempotency keys because retries happen at every layer.
  • Choose synchronous, streaming, background, scheduled or event-driven execution by run length.
  • Chaos-test resumption and target zero duplicate actions.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An approval webhook fires twice and the customer is refunded twice. What was missing?
  2. A research agent takes 8 minutes per run. What execution model fits best?
  3. When should an agent checkpoint its state?

Put it into practice

Add a runs table and the once() helper to your lesson 3 agent. Kill the process mid-run and confirm it resumes without repeating a side effect.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.