Computer-Use and Browser Agents: AI That Operates SoftwareReliability engineering for agents · Lesson 6 of 16

Task design, checkpoints and verification

Article · 8 min · 8 min lecture

Video lecture

Task design, checkpoints and verification

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Task design and verification

  • Why agents really fail
  • The five-part task template
  • Checkpoints
  • Verification that means something

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Reliability starts before the first click

Most agent failures are not model failures. They are task design failures: vague goals, no definition of done, no intermediate checks, and success judged by the agent's own claim. Long tasks compound error. If each step succeeds with high probability, a task with many steps can still fail often, because every step is another chance to go wrong. The engineering response is to make tasks shorter, checkable and restartable.

Write tasks like a test specification

A good agent task has five parts:

  1. Goal: one sentence, outcome-focused. "Record the listed price and availability for each product URL."
  2. Inputs: exactly what the agent receives (URL list, spreadsheet, account).
  3. Constraints: allowed domains, allowed actions, forbidden actions ("do not submit forms", "do not change settings").
  4. Definition of done: the structured output and its schema.
  5. Escalation rule: what to do when stuck, uncertain or asked for something sensitive ("return status NEEDS_HUMAN with the reason").
GOAL: For each URL in the input, record the product name, listed price (with currency)
and whether an "Add to cart" button is enabled.
INPUT: 25 product URLs on shop.example.pk (list below).
CONSTRAINTS: Stay on shop.example.pk. Do not click "Add to cart". Do not log in.
Dismiss cookie banners by choosing the minimum/necessary option.
DONE WHEN: You return a JSON array; one object per URL with keys
url, name, price, currency, add_to_cart_enabled, evidence_note.
IF STUCK: After 3 failed attempts on one URL, set status "NEEDS_HUMAN" and move on.
Never enter personal or payment data.

Decompose into checkpoints

Break long workflows into stages with checkpoints. At each checkpoint the harness saves state (URL, extracted data, screenshot) and verifies a condition before continuing. If a later stage fails, you resume from the last good checkpoint instead of starting over.

StageCheckpoint conditionEvidence saved
Open product pageURL matches expected pattern; HTTP 200Screenshot
Extract fieldsPrice parses as a number; currency presentJSON fragment
Verify button stateAccessibility node found with role buttonSnapshot excerpt

Verification: never trust "done"

An agent's final message is a claim. Verification turns claims into evidence. Four verification methods, from strongest to weakest:

  1. Ground-truth check. Query the system of record: the database, the CRM, the order API. If the agent claims it updated a CRM field, read the field back through the API.
  2. Deterministic assertion. Check the page state with code: the confirmation element exists, the URL changed to /thank-you, the table contains the expected row.
  3. Independent model verifier. A separate model call (ideally with a different prompt, sometimes a different model) examines the final screenshot and output and answers a narrow question: "Does this screenshot show a submitted form with a confirmation number? Answer yes/no and quote the number."
  4. Self-report. The agent says it succeeded. Useful only as a hint.

Use the strongest method available for each step. For read-only extraction, spot-check a sample against a human-verified answer key.

Hands-on: a verifier function

# verify.py
import os, json, base64
import anthropic

client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
VERIFIER_MODEL = os.environ.get("VERIFIER_MODEL", "claude-sonnet-5")

def verify_screenshot(png_bytes: bytes, question: str) -> dict:
    """Ask a narrow yes/no question about a final screenshot."""
    content = [
        {"type": "image", "source": {"type": "base64", "media_type": "image/png",
                                      "data": base64.b64encode(png_bytes).decode()}},
        {"type": "text", "text": (
            "You are a strict verifier. Answer only from what is visible. "
            f"Question: {question}\n"
            'Return JSON: {"answer": "yes"|"no"|"unclear", "evidence": "<quote what you see>"}')},
    ]
    try:
        msg = client.messages.create(model=VERIFIER_MODEL, max_tokens=300,
                                     messages=[{"role": "user", "content": content}])
        text = "".join(b.text for b in msg.content if b.type == "text")
        return json.loads(text)
    except (anthropic.APIError, json.JSONDecodeError) as exc:
        return {"answer": "unclear", "evidence": f"verifier error: {exc}"}

Treat "unclear" as a failure that goes to a human. Asymmetric costs matter: a false "yes" (claiming success that did not happen) is usually worse than a false "no".

Make tasks idempotent and restartable

An idempotent step produces the same result if run twice. Reading is naturally idempotent; writing is not. Before any write action, check whether it has already happened ("Is this listing already updated?"). Store a run ID and step IDs so a restarted run skips completed steps.

Worked example: a US e-commerce price-match check

A US retailer checks competitor prices on 50 products weekly. The first version asked the agent to "find competitor prices" and trusted its summary. Accuracy was poor and nobody knew which numbers were wrong. Version two used the task template above, one URL per sub-task, a price-parsing assertion, a screenshot per product and a 10% human spot-check. Errors became visible, fixable and measurable, and the team could state an accuracy figure from its own spot-check data rather than a guess.

Pitfalls

  • One giant prompt for a 40-step process. Split it.
  • Verifying with the same model call that did the work. Use an independent check.
  • No escalation path, so the agent improvises when stuck.

How to measure success

Track first-pass success, verified success (after checks), false-success rate (claimed but failed verification) and escalation rate. False-success rate is the number that should worry you most.

Key takeaways

  • Most failures are task-design failures; write tasks with goal, inputs, constraints, definition of done and escalation rule.
  • Split long workflows into checkpointed stages that save evidence and can resume.
  • Verify with ground truth or deterministic assertions first; use an independent model verifier next; treat self-reports as hints.
  • Make write steps idempotent and track false-success rate closely.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An agent reports 'CRM record updated'. What is the strongest verification?
  2. Which element is missing from this task: 'Check competitor prices and summarize'?
  3. Why is a false 'yes' from a verifier usually worse than a false 'no'?

Put it into practice

Rewrite one of your agent tasks using the five-part template, split it into checkpointed stages, and choose a verification method for each stage.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.