---
title: "Retries, recovery and observability | Optimize All Academy"
description: "Failure is normal; unhandled failure is not Web environments are noisy: slow loads, cookie banners, A/B tests, session timeouts, rate limits and pop-ups…"
url: https://optimizeall.com/learn/computer-use-and-browser-agents/retries-recovery-and-observability
updated: 2026-10-05
---

Computer-Use and Browser Agents: AI That Operates Software · Reliability engineering for agents · lesson 7 of 16 · 7 min

# Retries, recovery and observability

## Failure is normal; unhandled failure is not

Web environments are noisy: slow loads, cookie banners, A/B tests, session timeouts, rate limits and pop-ups. A reliable agent system expects these and handles them at the right layer. This lesson covers three disciplines: **retries** (trying again sensibly), **recovery** (getting back to a known state) and **observability** (seeing what happened, cheaply).

## Classify failures before retrying

Not every failure deserves a retry.

| Failure class | Examples | Response |
|---|---|---|
| Transient infrastructure | Timeout, 5xx, API rate limit (429), overloaded model | Retry with exponential backoff and jitter |
| Environment noise | Cookie banner, newsletter pop-up, slow spinner | Handle deterministically (scripted dismissal), then continue |
| Agent confusion | Clicking the wrong element repeatedly, looping | Re-ground: fresh screenshot, restate goal, or reset to checkpoint |
| Hard blocker | Login wall, CAPTCHA, access denied, missing data | Stop and escalate; never try to bypass |
| Policy trigger | Payment page, sending a message, deleting | Pause for human approval |

Retrying a hard blocker wastes money; retrying a policy trigger can be dangerous.

## Retry patterns

- **Exponential backoff with jitter** for API and network calls: wait, then wait longer, with randomness to avoid thundering herds. Respect `retry-after` headers.
- **Bounded retries per step** (for example three), then escalate.
- **Loop detection.** If the last N actions repeat the same pattern, or the screen has not changed after several actions, stop the loop and re-ground or reset.

```python
# resilience.py
import random, time, hashlib

def backoff_call(fn, *, attempts=4, base=1.0, cap=30.0, retry_on=(TimeoutError,)):
    for i in range(attempts):
        try:
            return fn()
        except retry_on:
            if i == attempts - 1:
                raise
            time.sleep(min(cap, base * 2 ** i) * random.uniform(0.5, 1.5))

class LoopGuard:
    """Flags when the screen stops changing or actions repeat."""
    def __init__(self, window=4):
        self.window, self.frames, self.actions = window, [], []

    def observe(self, png_bytes: bytes, action: str) -> bool:
        self.frames.append(hashlib.sha256(png_bytes).hexdigest())
        self.actions.append(action)
        f, a = self.frames[-self.window:], self.actions[-self.window:]
        stuck_screen = len(f) == self.window and len(set(f)) == 1
        repeating = len(a) == self.window and len(set(a)) == 1
        return stuck_screen or repeating      # True means: intervene
```

Exact hash matching is crude (a blinking cursor changes pixels), so production systems use perceptual hashes or compare accessibility snapshots; the principle is the same.

## Recovery: return to a known state

When an agent is lost, more reasoning on a confused state rarely helps. Options, in order:

1. **Re-ground**: send a fresh screenshot plus a short restatement of the goal and progress so far.
2. **Reset to checkpoint**: navigate to the last checkpoint URL, restore extracted data, continue.
3. **Fresh context**: new browser context (clears cookies and odd state) and restart the sub-task.
4. **Escalate** with evidence.

Handle environment noise **outside the model** where you can. A small Playwright routine that dismisses known cookie banners (choosing the necessary-only option, to respect consent) is cheaper and more reliable than letting the model figure it out on every page.

## Observability: make every run explainable

For each run, record:

- Run ID, task version, model and tool versions, start and end time.
- Every step: action, inputs, result, latency, tokens, and a screenshot reference.
- Checkpoint results and verification outcomes.
- Final status: success, verified success, escalated, stopped by limit, error.
- Cost: input and output tokens and the monetary estimate.

Store screenshots in object storage with retention rules (they can contain personal data). Standard tracing tools (OpenTelemetry-based tracing, or LLM observability platforms) let you view runs as traces with nested spans. A **session replay** (screenshot filmstrip with the actions overlaid) is the fastest way to debug a failed run and to show a client or auditor what happened.

## Worked example: an Abu Dhabi government-services helper

A business-setup consultancy in Abu Dhabi uses an agent to pre-check that client applications on a licensing portal have the right documents uploaded (read-only). The portal times out often. Their harness retries page loads with backoff, detects the session-expired page deterministically and hands it to a human to log in again (the agent never handles the password), and flags any application where the loop guard fires. Every run produces a filmstrip replay attached to the client file.

## Pitfalls

- Retrying everything the same way, including CAPTCHAs and access-denied pages.
- Logging too little (no replay) or too much (full-resolution screenshots of personal data kept forever).
- Letting the model burn steps on cookie banners.

## How to measure success

Track retries per run, loop-guard triggers, escalations by failure class, mean time to diagnose a failed run, and cost per successful task. A falling time-to-diagnose shows your observability works.

## Video lecture: Retries, recovery and observability

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. Retries, recovery and observability
2. Five failure classes
3. Why it matters
4. The delivery driver
5. Simple example: five overnight failures
6. Retry patterns
7. Recovery ladder
8. What to record
9. Privacy in logs
10. Worked example: licensing pre-checks
11. The five-minute test
12. Rate limits and politeness
13. Three mistakes
14. Try this now
15. Recap

## Lecture transcript

### Retries, recovery and observability

Web pages are messy. They load slowly, throw cookie banners at you, run A B tests, and time out your session at the worst moment. A reliable agent system expects all of that. In this lesson you will learn how to classify failures, retry the right ones, recover when an agent gets lost, and record every run so you can explain exactly what happened.

### Five failure classes

Start by classifying. Transient infrastructure failures, like timeouts, server errors or rate limits, deserve a retry with backoff. Environment noise, cookie banners and pop-ups, should be handled by a small script. Agent confusion, clicking the wrong thing repeatedly, needs re-grounding. Hard blockers, a login wall, a CAPTCHA, access denied, mean stop and escalate. And policy triggers, a payment page, sending a message, deleting something, mean pause for human approval. Retrying a hard blocker wastes money. Retrying a policy trigger can be dangerous.

### Why it matters

Why does this matter? Because an agent that fails silently, or fails expensively, will lose the trust of everyone around it, however clever it is. Reliability is what turns a demo into a tool people depend on. And observability is what lets you improve: without logs and replays, every failure is a mystery, and you end up guessing. With them, a new team member can diagnose a problem in minutes and you can show clients exactly what happened.

### The delivery driver

Here's an analogy. Think of a delivery driver. If traffic is heavy, they wait and try again, that's a transient failure. If the road has roadworks every day, the company updates the route, that's environment noise handled once. If the driver gets lost, they pull over, check the map and restart from the last known street, that's recovery. And if the building has a locked gate, they don't climb over it, they call the customer. Every class of failure has its own right response.

### Simple example: five overnight failures

A simple example of failure classification. Your agent reports five failures overnight. One is a timeout, retry with backoff. One is a cookie banner blocking the page, add a scripted dismissal so it never happens again. One is the agent clicking the wrong tab three times, re-ground it or reset to the checkpoint. One is a login wall, escalate to a human. And one is a page asking the agent to confirm a subscription, pause for approval. Five failures, five different responses, and none of them is just try again.

### Retry patterns

For transient failures, use exponential backoff with jitter. Wait a second, then two, then four, with a little randomness so many workers don't retry at the same instant. Respect retry after headers. Cap retries per step, three is a sensible default, then escalate. And add a loop guard. If the screen hasn't changed across several actions, or the agent keeps repeating the same action, stop and intervene. The lesson text includes a small loop guard you can adapt.

### Recovery ladder

When an agent is lost, more thinking on a confused screen rarely helps. Recover in order. Re-ground: send a fresh screenshot with a short restatement of the goal and progress. Reset to the last checkpoint. Start a fresh browser context, which clears odd cookies and state. And if none of that works, escalate with evidence. Meanwhile, take known noise away from the model entirely. A few lines of Playwright that dismiss common cookie banners, choosing the necessary-only option to respect consent, save steps on every single page.

### What to record

Now observability. For every run, record the run ID, task version, model and tool versions, and timing. For every step, the action, its inputs, the result, latency, tokens and a screenshot reference. Record checkpoint and verification results, the final status, and cost. Standard tracing tools let you view each run as a tree of spans. And a session replay, a filmstrip of screenshots with actions overlaid, is the fastest way to debug a failure or to show a client exactly what happened.

### Privacy in logs

One caution. Screenshots can contain personal data, names, emails, account numbers. Store them in controlled storage with retention rules, and don't keep full-resolution images forever. Observability and privacy have to be designed together, especially under GDPR, UK GDPR and the data protection laws in the Gulf.

### Worked example: licensing pre-checks

Here's how it looks in practice. A business-setup consultancy in Abu Dhabi pre-checks client applications on a licensing portal, read-only. The portal times out a lot, so page loads retry with backoff. A session-expired page is detected in code and handed to a human to log back in, the agent never touches the password. Any application where the loop guard fires is flagged. And every run produces a filmstrip replay attached to the client file.

### The five-minute test

How do you know your observability is good enough? Try this test. Pick a failed run from last week and time how long it takes someone new to explain what went wrong. If they can open the replay, see the step where the agent went off track, read the action and the screenshot, and name the failure class in a few minutes, you're in good shape. If they have to rerun it and hope it fails the same way, you need better logging.

### Rate limits and politeness

Rate limits deserve special care. Model APIs and websites both limit how fast you can go. When you get a rate-limit response, read the retry after header if there is one and wait at least that long. Spread scheduled runs out rather than starting everything at nine on Monday morning. And be a polite visitor to other people's websites: limit concurrent pages per domain, so your audit never looks like an attack or slows a small business's site.

### Three mistakes

Three common mistakes. First, retrying everything the same way, including CAPTCHAs and access-denied pages, which wastes money and can breach terms. Second, logging too little, so nobody can explain a failure, or logging too much, like full-resolution screenshots of personal data kept forever. Third, letting the model burn steps on cookie banners and pop-ups that a few lines of deterministic code could handle once and for all.

### Try this now

Try this now. Collect your last ten failed agent runs, or failed automated test runs if you don't have agents yet. For each, write down the failure class: transient, environment noise, agent confusion, hard blocker or policy trigger. Then write which layer should have handled it: a retry, a scripted fix, a recovery step, an escalation or an approval gate. Pick the most common class and implement one improvement for it this week, then compare failure counts next week.

### Recap

Recap. Classify failures into five classes and respond to each differently. Retry transient errors with backoff, cap retries, and guard against loops. Recover by re-grounding, resetting, or starting fresh, then escalate. Log every step and keep replays, with sensible retention. Your next step: take your last ten failed runs, label each with its failure class, and decide which layer should have handled it. Next, how to evaluate an agent properly.

## Key takeaways

- Classify failures: transient, environment noise, agent confusion, hard blocker, policy trigger; respond differently to each.
- Use bounded retries with exponential backoff and jitter, and detect loops when screens or actions repeat.
- Recover by re-grounding, resetting to a checkpoint or starting a fresh context, then escalate with evidence.
- Log every step with screenshots, tokens and cost; session replays make failures explainable, but set retention limits.

## Try it

Add a failure-class column to your last ten agent runs (or test runs). Which class was most common, and which layer should handle it?

- [Previous: Task design, checkpoints and verification](https://optimizeall.com/learn/computer-use-and-browser-agents/task-design-checkpoints-and-verification)
- [Next: Evaluating browser agents: test sets, metrics and benchmarks](https://optimizeall.com/learn/computer-use-and-browser-agents/evaluating-browser-agents)
- [All lessons of Computer-Use and Browser Agents: AI That Operates Software](https://optimizeall.com/learn/computer-use-and-browser-agents)
