Computer-Use and Browser Agents: AI That Operates SoftwareReliability engineering for agents · Lesson 7 of 16

Retries, recovery and observability

Article · 7 min · 8 min lecture

Video lecture

Retries, recovery and observability

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Retries, recovery and observability

  • Classify failures
  • Retry sensibly
  • Recover to a known state
  • Make every run explainable

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Failure is normal; unhandled failure is not

Web environments are noisy: slow loads, cookie banners, A/B tests, session timeouts, rate limits and pop-ups. A reliable agent system expects these and handles them at the right layer. This lesson covers three disciplines: retries (trying again sensibly), recovery (getting back to a known state) and observability (seeing what happened, cheaply).

Classify failures before retrying

Not every failure deserves a retry.

Failure classExamplesResponse
Transient infrastructureTimeout, 5xx, API rate limit (429), overloaded modelRetry with exponential backoff and jitter
Environment noiseCookie banner, newsletter pop-up, slow spinnerHandle deterministically (scripted dismissal), then continue
Agent confusionClicking the wrong element repeatedly, loopingRe-ground: fresh screenshot, restate goal, or reset to checkpoint
Hard blockerLogin wall, CAPTCHA, access denied, missing dataStop and escalate; never try to bypass
Policy triggerPayment page, sending a message, deletingPause for human approval

Retrying a hard blocker wastes money; retrying a policy trigger can be dangerous.

Retry patterns

  • Exponential backoff with jitter for API and network calls: wait, then wait longer, with randomness to avoid thundering herds. Respect retry-after headers.
  • Bounded retries per step (for example three), then escalate.
  • Loop detection. If the last N actions repeat the same pattern, or the screen has not changed after several actions, stop the loop and re-ground or reset.
# resilience.py
import random, time, hashlib

def backoff_call(fn, *, attempts=4, base=1.0, cap=30.0, retry_on=(TimeoutError,)):
    for i in range(attempts):
        try:
            return fn()
        except retry_on:
            if i == attempts - 1:
                raise
            time.sleep(min(cap, base * 2 ** i) * random.uniform(0.5, 1.5))

class LoopGuard:
    """Flags when the screen stops changing or actions repeat."""
    def __init__(self, window=4):
        self.window, self.frames, self.actions = window, [], []

    def observe(self, png_bytes: bytes, action: str) -> bool:
        self.frames.append(hashlib.sha256(png_bytes).hexdigest())
        self.actions.append(action)
        f, a = self.frames[-self.window:], self.actions[-self.window:]
        stuck_screen = len(f) == self.window and len(set(f)) == 1
        repeating = len(a) == self.window and len(set(a)) == 1
        return stuck_screen or repeating      # True means: intervene

Exact hash matching is crude (a blinking cursor changes pixels), so production systems use perceptual hashes or compare accessibility snapshots; the principle is the same.

Recovery: return to a known state

When an agent is lost, more reasoning on a confused state rarely helps. Options, in order:

  1. Re-ground: send a fresh screenshot plus a short restatement of the goal and progress so far.
  2. Reset to checkpoint: navigate to the last checkpoint URL, restore extracted data, continue.
  3. Fresh context: new browser context (clears cookies and odd state) and restart the sub-task.
  4. Escalate with evidence.

Handle environment noise outside the model where you can. A small Playwright routine that dismisses known cookie banners (choosing the necessary-only option, to respect consent) is cheaper and more reliable than letting the model figure it out on every page.

Observability: make every run explainable

For each run, record:

  • Run ID, task version, model and tool versions, start and end time.
  • Every step: action, inputs, result, latency, tokens, and a screenshot reference.
  • Checkpoint results and verification outcomes.
  • Final status: success, verified success, escalated, stopped by limit, error.
  • Cost: input and output tokens and the monetary estimate.

Store screenshots in object storage with retention rules (they can contain personal data). Standard tracing tools (OpenTelemetry-based tracing, or LLM observability platforms) let you view runs as traces with nested spans. A session replay (screenshot filmstrip with the actions overlaid) is the fastest way to debug a failed run and to show a client or auditor what happened.

Worked example: an Abu Dhabi government-services helper

A business-setup consultancy in Abu Dhabi uses an agent to pre-check that client applications on a licensing portal have the right documents uploaded (read-only). The portal times out often. Their harness retries page loads with backoff, detects the session-expired page deterministically and hands it to a human to log in again (the agent never handles the password), and flags any application where the loop guard fires. Every run produces a filmstrip replay attached to the client file.

Pitfalls

  • Retrying everything the same way, including CAPTCHAs and access-denied pages.
  • Logging too little (no replay) or too much (full-resolution screenshots of personal data kept forever).
  • Letting the model burn steps on cookie banners.

How to measure success

Track retries per run, loop-guard triggers, escalations by failure class, mean time to diagnose a failed run, and cost per successful task. A falling time-to-diagnose shows your observability works.

Key takeaways

  • Classify failures: transient, environment noise, agent confusion, hard blocker, policy trigger; respond differently to each.
  • Use bounded retries with exponential backoff and jitter, and detect loops when screens or actions repeat.
  • Recover by re-grounding, resetting to a checkpoint or starting a fresh context, then escalate with evidence.
  • Log every step with screenshots, tokens and cost; session replays make failures explainable, but set retention limits.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An agent hits a CAPTCHA on a supplier site. What should the harness do?
  2. The last four screenshots are identical and the agent keeps clicking the same spot. What is the best first response?
  3. Why handle cookie banners with a scripted routine?

Put it into practice

Add a failure-class column to your last ten agent runs (or test runs). Which class was most common, and which layer should handle it?

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.