Computer-Use and Browser Agents: AI That Operates SoftwareReliability engineering for agents · Lesson 6 of 16
Task design, checkpoints and verification
Video lecture
Task design, checkpoints and verification
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Task design and verification
Here is an uncomfortable truth. When an agent fails, it is usually not because the model was not smart enough. It is because the task was vague, nobody defined done, and success was judged by the agent's own claim. In this lesson you will learn to write agent tasks like test specifications, add checkpoints, and verify results so that done actually means done.
0:27 Why long tasks fail
Why do long tasks fail so often? Every step is another chance to go wrong. Even if each step usually works, string enough of them together and the odds of a clean run fall quickly. The engineering answer is not a smarter model. It is shorter tasks, checkable steps, and the ability to restart from the last good point.
0:53 Why it matters
Why does this matter? Because the cost of an agent isn't just tokens. It's the human time spent checking, fixing and explaining its work. A vague task produces outputs that take ages to check and are hard to trust. A well-specified task, with checkpoints and verification built in, produces outputs you can check in seconds and defend to a client or manager. Good task design is the cheapest reliability improvement available.
1:24 The shopping list
Here's an analogy. Think about sending a new assistant to the shops. Buy food for the weekend will get you something, but probably not what you wanted. A shopping list with brands, quantities, a budget, a rule like if they're out of oat milk call me, and please bring the receipt, gets you exactly what you need, and proof that it happened. Agent tasks work the same way. The receipt is your verification.
1:56 Simple example: 25 pages
A simple example of checkpoints. Suppose the agent must collect prices from twenty-five product pages. Without checkpoints, a timeout on page nineteen means starting again and paying twice. With checkpoints, after each page the harness saves the URL, the extracted price and a screenshot, and checks that the price parses as a number. When page nineteen fails, the run resumes at page nineteen, and pages one to eighteen are already done and verified.
2:28 The five-part task
Write every task in five parts. A goal, one outcome-focused sentence. Inputs, exactly what the agent gets. Constraints: which domains, which actions are allowed, and which are forbidden. A definition of done: the exact structured output you expect. And an escalation rule: what to do when stuck, uncertain or asked for something sensitive. For example: after three failed attempts on one page, mark it needs human and move on. That single rule stops agents improvising in the places where improvisation hurts most.
3:04 Checkpoints
Next, checkpoints. Split the workflow into stages. At the end of each stage the harness saves state, the URL, the extracted data, a screenshot, and checks a condition before moving on. Did the page return successfully? Does the price parse as a number? Is there a button with the right role? If a later stage fails, you resume from the last good checkpoint instead of starting over and paying twice.
3:34 Verification ladder
Now verification. An agent's final message is a claim, not evidence. Rank your options. Strongest is ground truth: read the record back from the database or CRM API. Next, deterministic assertions in code: the confirmation element exists, the URL changed. Then an independent verifier, a separate model call that looks at the final screenshot and answers one narrow question, quoting what it sees. Weakest is the agent's own report. Use the strongest method you can for each step.
4:08 Two habits
Two more habits. First, treat unclear as a failure that goes to a human. A verifier that wrongly says yes hides a problem until it reaches a customer. A verifier that wrongly says no just costs a quick human check. Second, make write steps idempotent. Before the agent changes anything, check whether the change already happened, and track run and step IDs so a restarted run skips completed work.
4:38 Worked example: price checks
A US retailer learned this the hard way. Version one asked an agent to find competitor prices and trusted the summary. Nobody knew which numbers were wrong. Version two used the five-part task, one product per sub-task, a price-parsing check, a screenshot per product, and a human spot-check of one in ten. Errors became visible and fixable, and the team could quote accuracy from its own data rather than a hunch.
5:09 Worked task spec
Let's write one together, quickly. Goal: record the listed price and availability for each product URL. Inputs: twenty-five URLs on one shop domain. Constraints: stay on that domain, don't log in, don't click add to cart, choose necessary-only cookies. Done when: a JSON list with URL, name, price, currency and availability. If stuck: after three failed attempts on one URL, mark it needs human and move on. Five short sections, and suddenly the agent's job, and your test, are both crystal clear.
5:45 Independent verifiers
A note on independent verifiers. When you use a second model call to check the result, keep it narrow and strict. Give it the final screenshot and one yes or no question, and ask it to quote what it sees as evidence. Don't show it the agent's reasoning, which can bias it toward agreeing. And where you can, use a different prompt or even a different model from the one that did the work. Independence is what makes a verifier worth having.
6:21 Three mistakes
Three common mistakes in task design. First, one giant prompt for a forty-step process. Split it into sub-tasks with their own checks. Second, verifying with the same model call that did the work. It tends to agree with itself. Use ground truth, code assertions or an independent verifier. Third, no escalation path. When the agent gets stuck and has no rule for what to do, it improvises, and improvisation is where damage happens.
6:53 Try this now
Try this now. Take one agent task you're using or planning, and rewrite it using the five-part template: goal, inputs, constraints, definition of done and escalation rule. Then split it into stages, and for each stage write one checkpoint condition and the evidence you'll save. Finally, pick a verification method for the final result, choosing the strongest one available from the ladder: ground truth, deterministic assertion, independent verifier or, only as a last resort, self-report.
7:26 Recap
Recap. Write tasks with a goal, inputs, constraints, definition of done and an escalation rule. Checkpoint long workflows. Verify with ground truth first, and treat unclear as failure. The metric to watch is false success, claimed but not real. Your next step: rewrite one of your own agent tasks with the template and choose a verification method for each stage. Next lesson: retries, recovery and observability.
Reliability starts before the first click
Most agent failures are not model failures. They are task design failures: vague goals, no definition of done, no intermediate checks, and success judged by the agent's own claim. Long tasks compound error. If each step succeeds with high probability, a task with many steps can still fail often, because every step is another chance to go wrong. The engineering response is to make tasks shorter, checkable and restartable.
Write tasks like a test specification
A good agent task has five parts:
- Goal: one sentence, outcome-focused. "Record the listed price and availability for each product URL."
- Inputs: exactly what the agent receives (URL list, spreadsheet, account).
- Constraints: allowed domains, allowed actions, forbidden actions ("do not submit forms", "do not change settings").
- Definition of done: the structured output and its schema.
- Escalation rule: what to do when stuck, uncertain or asked for something sensitive ("return status NEEDS_HUMAN with the reason").
GOAL: For each URL in the input, record the product name, listed price (with currency)
and whether an "Add to cart" button is enabled.
INPUT: 25 product URLs on shop.example.pk (list below).
CONSTRAINTS: Stay on shop.example.pk. Do not click "Add to cart". Do not log in.
Dismiss cookie banners by choosing the minimum/necessary option.
DONE WHEN: You return a JSON array; one object per URL with keys
url, name, price, currency, add_to_cart_enabled, evidence_note.
IF STUCK: After 3 failed attempts on one URL, set status "NEEDS_HUMAN" and move on.
Never enter personal or payment data.Decompose into checkpoints
Break long workflows into stages with checkpoints. At each checkpoint the harness saves state (URL, extracted data, screenshot) and verifies a condition before continuing. If a later stage fails, you resume from the last good checkpoint instead of starting over.
| Stage | Checkpoint condition | Evidence saved |
|---|---|---|
| Open product page | URL matches expected pattern; HTTP 200 | Screenshot |
| Extract fields | Price parses as a number; currency present | JSON fragment |
| Verify button state | Accessibility node found with role button | Snapshot excerpt |
Verification: never trust "done"
An agent's final message is a claim. Verification turns claims into evidence. Four verification methods, from strongest to weakest:
- Ground-truth check. Query the system of record: the database, the CRM, the order API. If the agent claims it updated a CRM field, read the field back through the API.
- Deterministic assertion. Check the page state with code: the confirmation element exists, the URL changed to /thank-you, the table contains the expected row.
- Independent model verifier. A separate model call (ideally with a different prompt, sometimes a different model) examines the final screenshot and output and answers a narrow question: "Does this screenshot show a submitted form with a confirmation number? Answer yes/no and quote the number."
- Self-report. The agent says it succeeded. Useful only as a hint.
Use the strongest method available for each step. For read-only extraction, spot-check a sample against a human-verified answer key.
Hands-on: a verifier function
# verify.py
import os, json, base64
import anthropic
client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
VERIFIER_MODEL = os.environ.get("VERIFIER_MODEL", "claude-sonnet-5")
def verify_screenshot(png_bytes: bytes, question: str) -> dict:
"""Ask a narrow yes/no question about a final screenshot."""
content = [
{"type": "image", "source": {"type": "base64", "media_type": "image/png",
"data": base64.b64encode(png_bytes).decode()}},
{"type": "text", "text": (
"You are a strict verifier. Answer only from what is visible. "
f"Question: {question}\n"
'Return JSON: {"answer": "yes"|"no"|"unclear", "evidence": "<quote what you see>"}')},
]
try:
msg = client.messages.create(model=VERIFIER_MODEL, max_tokens=300,
messages=[{"role": "user", "content": content}])
text = "".join(b.text for b in msg.content if b.type == "text")
return json.loads(text)
except (anthropic.APIError, json.JSONDecodeError) as exc:
return {"answer": "unclear", "evidence": f"verifier error: {exc}"}Treat "unclear" as a failure that goes to a human. Asymmetric costs matter: a false "yes" (claiming success that did not happen) is usually worse than a false "no".
Make tasks idempotent and restartable
An idempotent step produces the same result if run twice. Reading is naturally idempotent; writing is not. Before any write action, check whether it has already happened ("Is this listing already updated?"). Store a run ID and step IDs so a restarted run skips completed steps.
Worked example: a US e-commerce price-match check
A US retailer checks competitor prices on 50 products weekly. The first version asked the agent to "find competitor prices" and trusted its summary. Accuracy was poor and nobody knew which numbers were wrong. Version two used the task template above, one URL per sub-task, a price-parsing assertion, a screenshot per product and a 10% human spot-check. Errors became visible, fixable and measurable, and the team could state an accuracy figure from its own spot-check data rather than a guess.
Pitfalls
- One giant prompt for a 40-step process. Split it.
- Verifying with the same model call that did the work. Use an independent check.
- No escalation path, so the agent improvises when stuck.
How to measure success
Track first-pass success, verified success (after checks), false-success rate (claimed but failed verification) and escalation rate. False-success rate is the number that should worry you most.
Key takeaways
- Most failures are task-design failures; write tasks with goal, inputs, constraints, definition of done and escalation rule.
- Split long workflows into checkpointed stages that save evidence and can resume.
- Verify with ground truth or deterministic assertions first; use an independent model verifier next; treat self-reports as hints.
- Make write steps idempotent and track false-success rate closely.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Rewrite one of your agent tasks using the five-part template, split it into checkpointed stages, and choose a verification method for each stage.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.