---
title: "State, durability and long-running agents"
description: "Why durability becomes the problem A demo agent runs in one process for 30 seconds. A production agent may run for minutes or hours, wait for human…"
url: https://optimizeall.com/learn/ai-agents-engineering/state-durability-long-running
updated: 2026-10-05
---

Building Production AI Agents · Memory, retrieval and durable state · lesson 9 of 18 · 15 min

# State, durability and long-running agents

## Why durability becomes the problem

A demo agent runs in one process for 30 seconds. A production agent may run for minutes or hours, wait for human approval overnight, survive deploys and crash mid-step. Without durable state, a crash means lost work, duplicated side effects (two refunds, two emails), or a confused user. **Durability** is the property that an agent's progress survives failures and can resume exactly where it stopped.

## What counts as agent state

- The message history (including tool calls and results).
- The plan or todo list and which steps are complete.
- Structured facts (budgets, IDs, approvals granted).
- Pending actions awaiting approval.
- Budgets consumed (steps, tokens, money).
- References to external artifacts (files, drafts, tickets).

Store it outside the process: a database row per run, keyed by `run_id`, updated after every step.

## Checkpointing

After every model call and tool execution, persist a **checkpoint**. On restart, load the latest checkpoint and continue the loop. Frameworks do this for you: LangGraph compiles graphs with a **checkpointer** (in-memory for dev; Postgres or others for production) and supports `interrupt()` for pauses; the OpenAI Agents SDK has **sessions** and serialisable run state for resuming after approvals; managed agent platforms keep sessions server-side. Durable-execution engines such as Temporal are also used to wrap agent steps as retriable activities.

## Idempotency: the rule that prevents double actions

Any tool with side effects must be safe to retry. Pattern:

1. Generate an **idempotency key** for each intended action (for example `run_id + step + tool_name`).
2. Pass it to the downstream API (many payment and messaging APIs accept one) or record it in your own table before executing.
3. If the key already exists with a result, return the stored result instead of acting again.

This matters because retries happen at every layer: SDK retries, queue redelivery, your own resume logic.

## Execution models for long runs

| Model | How it works | Use when |
|---|---|---|
| Synchronous request | API call waits for the agent to finish | Runs under ~30–60 s, interactive chat |
| Streaming | Stream progress events to the UI | Interactive but multi-step |
| Background job | Enqueue run; worker processes; client polls or gets webhook | Minutes-long research, batch work |
| Scheduled | Cron triggers runs | Daily reports, monitoring |
| Event-driven | Webhooks (new lead, new ticket) start runs | Ops automation |

Background jobs need a queue (e.g., a managed queue service or Redis-backed workers), a runs table, and status endpoints.

## Worked example: overnight competitor monitoring for a US DTC brand

Every night at 02:00 local time, an agent checks 30 competitor sites for price changes, drafts a summary and proposes price responses for approval.

- Scheduled trigger creates `run_id`, enqueues one job per competitor (fan-out) plus a final synthesis job.
- Each job checkpoints after every tool call; a crash at competitor 17 resumes at 17.
- Proposed price changes are stored as **pending actions** with idempotency keys; the pricing manager approves in the morning; execution happens once even if the approval webhook fires twice.
- Budgets: 12 tool calls per competitor, a nightly token cap, alert if exceeded.

## Hands-on: a resumable loop with a runs table

```python
import json, sqlite3, uuid

db = sqlite3.connect("runs.db")
db.execute("""CREATE TABLE IF NOT EXISTS runs(
  run_id TEXT PRIMARY KEY, status TEXT, step INTEGER, messages TEXT, updated REAL)""")
db.execute("""CREATE TABLE IF NOT EXISTS actions(
  idem_key TEXT PRIMARY KEY, result TEXT)""")

def save(run_id, status, step, messages):
    db.execute("REPLACE INTO runs VALUES (?,?,?,?,strftime('%s','now'))",
               (run_id, status, step, json.dumps(messages, default=str)))
    db.commit()

def load(run_id):
    row = db.execute("SELECT status, step, messages FROM runs WHERE run_id=?", (run_id,)).fetchone()
    return (row[0], row[1], json.loads(row[2])) if row else None

def once(idem_key: str, action):
    """Run a side-effecting action at most once per key."""
    row = db.execute("SELECT result FROM actions WHERE idem_key=?", (idem_key,)).fetchone()
    if row:
        return json.loads(row[0])
    result = action()
    db.execute("INSERT INTO actions VALUES (?,?)", (idem_key, json.dumps(result)))
    db.commit()
    return result

def start(goal: str) -> str:
    run_id = str(uuid.uuid4())
    save(run_id, "running", 0, [{"role": "user", "content": goal}])
    return run_id

# In the loop from lesson 3: after each step call save(run_id, "running", step, messages).
# Before side effects: once(f"{run_id}:{step}:{tool_name}", lambda: do_the_thing(**args)).
# On worker start: status, step, messages = load(run_id) and continue from there.
```

Note that SDK response objects must be serialized to plain dicts (for example with `.model_dump()` on Anthropic/OpenAI Pydantic models) before storing.

## Pitfalls

- Holding state only in memory in a web process that autoscaling can kill.
- Retrying side-effecting tools without idempotency keys.
- Very long synchronous HTTP requests that time out at load balancers.
- No status endpoint, so users refresh and start duplicate runs.

## Measuring success

Track resumption success after injected failures (kill a worker mid-run in staging), duplicate-action rate (target zero), run completion rate, and time-to-resume. Chaos-test durability before you trust it.

## Video lecture: State, durability and long-running agents

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. State and durability
2. Why it matters
3. What is agent state?
4. Checkpointing
5. Idempotency
6. Simple example: invoice reminders
7. Execution models
8. Example: nightly competitor monitor
9. Hands-on: resumable loop
10. Pitfalls and metrics
11. How long to keep run state?
12. Deeper: the nightly monitor (illustrative)
13. Watch me do it: resume after a crash
14. Try this now
15. Recap

## Lecture transcript

### State and durability

A demo agent runs for thirty seconds in one process. A real one might run for an hour, wait overnight for approval, survive a deploy, and crash halfway through a step. If it can't pick up where it left off, you get lost work, duplicate emails, or worse, two refunds. In this lesson you'll learn how to make agents durable: checkpoints, idempotency and the right execution model.

### Why it matters

Why does durability matter? Because the moment agents do real work, like sending messages, updating records or spending money, a crash in the middle is no longer just annoying. It's expensive or embarrassing. Think of a video game with save points. Without them, dying at level nine sends you back to level one. With them, you respawn at the last save. Checkpoints are save points for agents. And idempotency is the rule that the treasure chest you already opened doesn't give you the gold twice when you reload.

### What is agent state?

First, what is agent state? It's the message history, including tool calls and results. The plan or todo list and what's done. Structured facts like budgets and approved IDs. Pending actions waiting for approval. Budgets consumed so far. And links to external artifacts like drafts or tickets. All of this belongs outside the process, in a database row for each run, updated after every step.

### Checkpointing

Checkpointing means saving that state after every model call and every tool execution. If the worker dies, a new one loads the latest checkpoint and continues. Frameworks help. LangGraph compiles graphs with a checkpointer and supports interrupts for pauses. The OpenAI Agents SDK has sessions and serialisable run state for resuming after approvals. Managed agent platforms keep sessions on the server. Some teams wrap each step in a durable execution engine like Temporal. Whatever you choose, test it by killing a worker on purpose.

### Idempotency

Now the rule that prevents double actions: idempotency. Retries happen everywhere. The SDK retries, queues redeliver messages, and your resume logic replays steps. So every tool with side effects must be safe to repeat. Give each intended action a key, such as the run id plus the step plus the tool name. Before acting, check whether that key already has a result. If it does, return the stored result instead of acting again. Many payment and messaging APIs accept idempotency keys directly.

### Simple example: invoice reminders

A simple example. An agent sends a weekly invoice reminder email to each overdue client. The worker crashes after sending three of ten. Without durable state, the retry starts again from client one, and the first three get a second, slightly annoyed reminder. With a runs table and an idempotency key per client per week, the retry checks the key, sees that clients one to three are done, and continues from client four. Same code, same crash, but a very different customer experience.

### Execution models

Pick the execution model by how long runs take. Short interactive runs can be a synchronous request. Multi step but interactive: stream progress events to the interface. Minutes long research: a background job, where the client polls a status endpoint or gets a webhook. Daily reports: scheduled runs. And operational automation: event driven runs, triggered by things like a new lead or a new ticket.

### Example: nightly competitor monitor

Here's a worked example. A US direct to consumer brand runs a nightly competitor monitor at two in the morning. It creates a run and fans out one job per competitor, thirty in total, plus a synthesis job. Each job checkpoints after every tool call, so a crash at competitor seventeen resumes at seventeen. Proposed price changes are stored as pending actions with idempotency keys. The pricing manager approves in the morning, and the change happens exactly once, even if the approval webhook fires twice.

### Hands-on: resumable loop

The lesson code gives you a minimal version with SQLite. A runs table stores status, step and messages. An actions table stores results by idempotency key. The once helper runs a side effect at most once per key. After each loop step you save, and on restart you load and continue. One practical gotcha: SDK response objects must be converted to plain dictionaries before you store them, for example with model dump.

### Pitfalls and metrics

Common pitfalls: keeping state only in a web process that autoscaling can kill, retrying side effects without keys, long synchronous requests that time out at the load balancer, and no status endpoint, so users refresh and start duplicate runs. Measure resumption success after injected failures, duplicate action rate, which should be zero, completion rate and time to resume.

### How long to keep run state?

People often ask how long to keep run state. Keep active runs until they finish or expire, then keep a compact record for as long as you need it for audits and debugging, typically weeks or months, not forever. Store the final outcome, the key decisions, approvals and costs, and drop bulky intermediate tool outputs, especially if they contain personal data. Set a retention policy, automate cleanup, and make sure it matches what your privacy notice promises.

### Deeper: the nightly monitor (illustrative)

Let's deepen the US brand's nightly monitor. Thirty competitors, one job each, run at two in the morning. In the first month, illustrative again, workers crashed three times because a competitor site timed out repeatedly and the container ran out of memory. Each time, the run resumed from the last checkpoint and finished before the pricing manager started work. On one night the approval webhook fired twice because of a network retry, and the idempotency key meant the price changed exactly once. The team's favorite metric became time to resume after a crash, which stayed under a minute. Without checkpoints, each of those nights would have meant a partial report and a manual rerun.

### Watch me do it: resume after a crash

Watch me do it. Let's step through the resumable loop. Two tables: runs, with a run id, status, step, the serialized messages and an updated time, and actions, keyed by an idempotency key with the stored result. Save writes the current status, step and messages. Load returns them for a run id. The once helper is the heart of it: it looks up the key, and if a result exists it returns that result without acting; otherwise it runs the action, stores the result and returns it. Start creates a run id and saves the goal. In the agent loop, after every step I call save, and before any side effect I call once with a key built from the run id, the step and the tool name. Now I kill the process with control C after step three, restart with the same run id, load the state, and it continues at step four. The file write happened exactly once.

### Try this now

Try this now. Add the runs table and the once helper from the lesson to your own agent. Pick one tool with a side effect, even a fake one that writes to a file. Start a run, and kill the process halfway through with Control C. Restart it from the saved run id. Did it continue from the right step? Did the side effect happen exactly once? If either answer is no, you've just found a production bug in a safe place.

### Recap

To recap: durable agents persist state outside the process, checkpoint after every step, make side effects idempotent, and use the right execution model for their run length. Your next step: add a runs table and the once helper to your agent from lesson three. Kill the process mid run and confirm it resumes without repeating a side effect.

## Key takeaways

- Durability means an agent's progress survives crashes, deploys and overnight waits.
- Persist a checkpoint after every model call and tool execution, keyed by run ID.
- Every side-effecting tool needs idempotency keys because retries happen at every layer.
- Choose synchronous, streaming, background, scheduled or event-driven execution by run length.
- Chaos-test resumption and target zero duplicate actions.

## Try it

Add a runs table and the once() helper to your lesson 3 agent. Kill the process mid-run and confirm it resumes without repeating a side effect.

- [Previous: Agentic retrieval: giving agents knowledge they can trust](https://optimizeall.com/learn/ai-agents-engineering/retrieval-for-agents)
- [Next: Human-in-the-loop: approvals, escalation and review](https://optimizeall.com/learn/ai-agents-engineering/human-in-the-loop-approvals)
- [All lessons of Building Production AI Agents](https://optimizeall.com/learn/ai-agents-engineering)
