Computer-Use and Browser Agents: AI That Operates SoftwareCapstone: a safe website-audit agent · Lesson 16 of 16
Capstone: a safe browser agent that audits a website with approval gates
Video lecture
Capstone: a safe browser agent that audits a website with approval gates
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Capstone: a safe website-audit agent
This is where it all comes together. You're going to design a browser agent that audits a website and produces a report a client would pay for, with human approval before anything that affects the outside world. By the end of this lesson you'll know exactly how the pieces fit, and how to prove it's safe before its first real run.
0:27 The brief
Here's the brief. Audit a client's site for at least five things: broken links, titles and meta tags, structured data consistency, accessibility basics, consent behavior, whether the key form works, pricing consistency, and agent-friendliness. Three actions are gated, meaning a human must approve: submitting a test inquiry, creating tickets in the client's tracker, and emailing the report.
0:52 Why it matters
Why does this matter? Because a website audit is one of the most useful, most sellable and safest things a browser agent can do. Agencies can offer it as a service, in-house teams can run it before every campaign, and it exercises every skill in this course: hybrid automation, sandboxing, injection defenses, approvals, evaluation and reporting. If you can build this safely, you can build most practical browser agents.
1:22 The building inspector
Here's an analogy for the capstone. Think of a building inspector. They don't renovate your house. They walk through with a checklist, take photos, measure things, and write a report with findings, severity and recommendations. If they need to test something that affects you, like turning off the water supply, they ask first. Your audit agent is a building inspector for websites: observe, measure, evidence, report, and ask before touching anything that matters.
1:54 Simple example: one finding
A simple example of one finding, end to end. The crawler visits a product page and records the price shown, three thousand four hundred and ninety-nine rupees, and the structured data price, two thousand nine hundred and ninety-nine. The judge model confirms the mismatch and rates it high severity, because assistants and search engines may show the wrong price. The report entry includes both values, a screenshot, the business impact in one sentence, and the fix: update the template's structured data. Evidence, impact, fix.
2:31 Architecture
The architecture has four parts. An orchestrator enforces limits, keeps the register and handles approvals. A deterministic Playwright crawler visits allowlisted pages and records status codes, titles, structured data, accessibility snapshots and network requests. A judge model reads that content and returns JSON verdicts, but never acts. And a computer-use sub-agent is used only for interactive checks, like the consent flow and the form, with restricted actions and submit buttons blocked unless approved.
3:03 Two interactive checks
Two checks deserve detail. For consent, the sub-agent clicks the reject option, and you compare network requests before and after. If tracking hosts fire before consent, that's a finding with real legal weight. For the inquiry form, the agent fills clearly marked test data, checks validation, and stops before submitting. It requests approval, a human says yes or no, and if yes, the result is verified with a screenshot check of the thank-you state.
3:35 Evaluate first
Before any client sees it, evaluate it. Build a golden site you control with seeded problems: a broken link, a missing meta description, a price mismatch in the structured data, an unlabeled field, a tag firing before consent, and a hidden prompt-injection instruction. Run the agent three times. You need every seeded issue found, zero actions on the injected instruction, and zero unapproved submissions. Record cost, duration and steps.
4:05 The report
Then the report. For every finding: the page, the check, the severity, the evidence, why it matters in plain business language, the recommended fix, and your confidence. Open with a two-minute executive summary. Close with what you couldn't check, like login areas, blocked pages and escalations. Honest limits build trust, and they stop a client assuming you checked something you didn't.
4:32 Governance and first run
Governance wraps it up. Register the workflow as tier two because it has gated submissions. Agree scope with the client in writing: which domains, what test data, whether test inquiries are acceptable, and who approves. Use a per-client service identity, state your retention period, and tell the client which AI services process their content. Here's a first run: a UK and UAE agency audits a forty-page site, finds analytics firing before consent and a phone field rejecting plus nine seven one numbers, gets approval for one test inquiry, verifies it, and the client books a fix sprint.
5:14 Write for the reader
Let's talk about the report's language. Your reader is often a marketing manager or business owner, not an engineer. So instead of saying analytics script loads before consent state is set, say your analytics starts tracking visitors before they agree, which may breach privacy rules and damage trust. Then give the fix in one sentence for their developer. Findings that explain business impact get fixed. Findings written in jargon get parked.
5:45 Know your unit cost
Finally, cost and timing. Record what each audit costs in tokens and infrastructure, how long it runs, and how many human minutes it takes to review. After a few audits you'll know your real unit cost, which lets you price the service sensibly, whether you're an agency selling audits or an in-house team justifying the time. And keep the golden site. Every time you change the agent, run it again before the next client audit.
6:18 Three mistakes
Three common mistakes on capstone projects. First, letting the computer-use sub-agent do everything, when the deterministic crawler could gather most evidence faster and cheaper. Second, testing only on friendly sites, and never including an injection test page. Third, reports written in jargon that clients can't act on. Put business impact first, keep fixes to one sentence, and be honest about what you couldn't check.
6:46 Try this now
Try this now. Build a small golden site you control with six seeded issues: a broken link, a missing meta description, a price mismatch in the structured data, an unlabeled form field, an analytics tag that fires before consent, and a hidden prompt-injection instruction. Run your audit agent against it three times. You're aiming for all six issues found, zero actions taken on the injected instruction, and zero form submissions without approval. Record steps, time and cost, then write your first report.
7:22 Course complete
You've finished the course. You can explain how computer-use agents see and act, build a harness with limits and logging, combine scripts with models, evaluate agents properly, defend them against injection, keep secrets and money under human control, understand agent protocols, make your own site agent-friendly, and run a governed program. Your final step: build this capstone against a golden site with at least six seeded issues, then take the final assessment. Well done.
The brief
Build a browser agent that audits a website for a client and produces an evidence-backed report, with human approval gates before any action that has an external effect. It must apply everything from this course: hybrid automation, sandboxing, allowlists, injection defenses, verification, observability, evaluation and governance.
Scope of the audit (pick at least five):
- Broken links and error pages (4xx/5xx).
- Page titles, meta descriptions and canonical tags present and sensible.
- Structured data present and consistent with visible content.
- Accessibility basics: unnamed controls, unlabeled fields, keyboard trap on key task.
- Consent banner offers a clear reject option and tags wait for consent (observe network requests).
- Key conversion task completable (for example, the inquiry form fills correctly) without submitting unless approved.
- Offer or pricing consistency against a provided source of truth.
- Agent-friendliness: facts in text, llms.txt present (optional), policies readable.
Gated actions (each needs explicit human approval): submitting a test inquiry, creating tickets in the client's issue tracker, emailing the report to the client.
Architecture
+-------------------+ approval requests +------------------+
task.yaml | Orchestrator |----------------------------->| Review UI/chat |
--------->| (limits, register, |<-----------------------------| (human) |
| golden-set gate) | decisions +------------------+
+---------+---------+
|
+---------------+----------------+
| |
+-----v------+ +------v-------+
| Crawler | pages, snapshots| Judge model | JSON verdicts only
| (Playwright|----------------->| (no actions) |
| sandbox) | +--------------+
+-----+------+
| only when a check needs interaction
+-----v-------------------+
| Computer-use sub-agent | restricted actions, allowlist, loop guard
+-------------------------+
|
evidence store (screenshots, HAR excerpts, logs; 30-day retention)Design choices to justify in your write-up:
- Deterministic crawler first: Playwright visits allowlisted URLs from the sitemap (capped), records status codes, titles, meta tags, JSON-LD, accessibility snapshots and network requests before and after consent.
- Judge model reads, never acts: evaluates content-level checks (offer consistency, meta quality, structured-data consistency) with untrusted-data instructions and JSON output.
- Computer-use sub-agent only for interactive checks (form completion, consent flow), with typing restricted to test data, submit buttons blocked unless approved.
- Approval gates for the three gated actions, with specific, evidence-backed requests.
Hands-on: the orchestrator core
# audit_agent.py (core flow; wire in modules from earlier lessons)
import json, os, time, uuid
from urllib.parse import urlparse
from playwright.sync_api import sync_playwright
from guards import check_navigation # module 4
from approval import request_approval, wait_for_decision
from verify import verify_screenshot # module 3
from judge import judge_page # module 2 pattern, JSON verdicts
RUN_ID = str(uuid.uuid4())[:8]
MAX_PAGES, BUDGET_USD = int(os.environ.get("MAX_PAGES", 50)), float(os.environ.get("BUDGET_USD", 10))
def crawl(start_urls, ctx):
findings, page = [], ctx.new_page()
requests_log = []
page.on("request", lambda r: requests_log.append(r.url))
for url in start_urls[:MAX_PAGES]:
try:
check_navigation(url)
except PermissionError as exc:
findings.append({"url": url, "severity": "info", "issue": str(exc)}); continue
requests_log.clear()
resp = page.goto(url, wait_until="load", timeout=30000)
shot = f"evidence/{RUN_ID}/{abs(hash(url))}.png"
page.screenshot(path=shot, full_page=True)
snapshot = page.locator("body").aria_snapshot()
jsonld = page.locator('script[type="application/ld+json"]').all_inner_texts()
record = {"url": url, "status": resp.status if resp else None, "title": page.title(),
"jsonld": jsonld, "pre_consent_requests": list(requests_log), "evidence": shot}
record["unnamed_buttons"] = sum(1 for ln in snapshot.splitlines() if ln.strip() in ("- button", "- button:"))
record["verdict"] = judge_page(page.locator("body").inner_text()[:15000], record)
findings.append(record)
return findings
def gated(action, domain, preview, evidence, reason, store):
req_id = request_approval(store, action=action, target_domain=domain, data_preview=preview,
screenshot_path=evidence, reason=reason)
approved, approver = wait_for_decision(store, req_id)
log_event("approval", {"action": action, "approved": approved, "approver": approver})
return approved
def log_event(kind, data):
with open(f"evidence/{RUN_ID}/events.jsonl", "a") as f:
f.write(json.dumps({"ts": time.time(), "run": RUN_ID, "kind": kind, **data}) + "\n")Complete the build by adding: the consent-flow check (the sub-agent clicks the reject option, then you compare network requests for tracking hosts before and after), the inquiry-form check (fill with clearly marked test data, verify field validation, stop before submit, request approval to submit, then verify the thank-you state with verify_screenshot), and the report generator (Markdown or HTML with severity, evidence links and fix recommendations).
The report
For each finding: page, check, severity (critical, high, medium, low), evidence (screenshot and data), why it matters (business impact in plain language), recommended fix, and confidence. Open with an executive summary a client can read in two minutes. Close with "what we could not check" (login areas, blocked pages, escalations), because honest limits build trust.
Evaluation before first client use
- Build a golden site: a small test site you control with seeded issues (a broken link, a missing meta description, a mismatched price in JSON-LD, an unlabeled field, a tag firing before consent, a hidden prompt-injection instruction).
- Run three times; require all seeded issues found, zero actions taken on the injected instruction, zero unapproved submissions.
- Record cost, duration and steps.
Governance and client communication
Register the workflow (tier 2, because of gated submissions), agree scope in writing with the client (domains, test data, whether test inquiries are acceptable, who approves), use a per-client service identity, and state retention. Tell the client which AI services process their site content.
Assessment rubric
| Criterion | Excellent |
|---|---|
| Safety | Allowlist enforced in two layers; all external effects gated; injection test passed |
| Reliability | Checkpoints, retries by failure class, loop guard, verification of gated actions |
| Evidence | Every finding has screenshot or data evidence and a confidence rating |
| Evaluation | Golden site with seeded issues; repeated runs; metrics reported |
| Usefulness | Clear executive summary, prioritized fixes, honest limits |
| Governance | Register entry, client scope agreement, retention and vendor disclosure |
Worked example: first client run
A UK-and-UAE agency runs the capstone agent on a new client's 40-page site. It finds a price mismatch between a product page and its JSON-LD, analytics firing before consent, three unnamed icon buttons in the header, and a contact form whose phone field rejects "+971" formatting. It requests approval to submit one test inquiry; the account manager approves; the agent verifies the thank-you page. The report leads with the consent issue (legal risk) and the phone-format bug (lost leads), and the client approves a fix sprint.
How to measure success
Seeded-issue recall on your golden site, false-positive rate on client sites (as judged by your team), human review minutes per audit, and client acceptance of recommendations.
Key takeaways
- Combine a deterministic crawler, a read-only judge model and a restricted computer-use sub-agent only for interactive checks.
- Gate every external effect (test submissions, tickets, emails) behind specific, evidence-backed approvals.
- Evaluate on a golden site with seeded issues and an injection test before any client use.
- Deliver evidence-backed findings with severity, business impact, fixes, confidence and honest limits.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Build the capstone agent against a golden site you control with at least six seeded issues, including one injection. Submit the report, run metrics and your register entry.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.