AI Free course · Certificate included
Computer-Use and Browser Agents: AI That Operates Software
Build, secure and govern AI agents that see screens, click, type and audit websites, with humans in control
- Advanced
- 4 h 53 min
- 16 lessons in 7 modules
- 2 h 13 min of video lectures
- Updated Sep 2026
About this course
AI agents can now operate software the way people do: reading screens, clicking, typing and navigating websites. This advanced course shows you how computer-use and browser agents really work, from screenshots, the DOM and accessibility trees to the agent loop in code, and how today's offerings from Anthropic, OpenAI and Google, AI browsers and Playwright-based tools compare. You will engineer reliability with task specs, checkpoints, verification, retries and golden-set evaluation, then lock agents down with sandboxes, least privilege, prompt-injection defenses, safe credential handling and human approval gates. You will learn how agent protocols such as MCP, A2A, ACP, UCP and AP2 are reshaping commerce, make your own website agent-friendly, and apply agents to marketing operations, QA and data entry. The capstone is a safe website-audit agent with evidence and approvals.
Tools you’ll use
- Claude API computer use
- OpenAI Responses API
- Gemini API
- Playwright
- Playwright MCP
- Model Context Protocol
- A2A protocol
- Docker
- Squid proxy
- Claude in Chrome
- schema.org
Skills
- Computer-use agents
- Browser automation
- Agent reliability engineering
- Prompt injection defense
- Agent sandboxing
- Agentic commerce
- Agent-friendly websites
- AI governance
What you’ll be able to do
- Explain how GUI agents perceive screens and act, and build a minimal computer-use harness with limits and logs
- Choose between APIs, scripts, hybrid automation and agents, and evaluate vendor tools with a durable checklist
- Design checkable agent tasks with checkpoints, verification, retries, observability and golden-set evaluation
- Sandbox agents with least privilege and defend against prompt injection on web pages
- Keep credentials and payments under human control with specific, risk-tiered approval gates
- Explain MCP, A2A, ACP, UCP and AP2 and make a website accurate and usable for AI agents
- Build and govern a safe browser agent that audits a website with evidence and approvals
Recommended first
Curriculum
Syllabus
- Modules
- 7
- Lessons
- 16
- Reading time
- 2 h
- Assessment questions
- 25
The perceive-reason-act loop, how agents see screens through pixels, the DOM and accessibility trees, and a minimal harness in code.
- From RPA to computer-use agents: what changedVideo lecture, 10′8 min
- How agents see: pixels, the DOM and accessibility treesVideo lecture, 9′8 min
- Writing the agent loop: a minimal computer-use harnessVideo lecture, 9′9 min
The 2026 landscape of computer-use APIs, agent products and AI browsers, a durable evaluation checklist, and hybrid patterns that combine Playwright with model reasoning.
- The computer-use landscape: APIs, agent products and AI browsersVideo lecture, 8′8 min
- Hybrid automation: Playwright scripts plus model reasoningVideo lecture, 9′7 min
Designing checkable tasks, checkpoints and verification, handling retries and recovery with full observability, and evaluating agents with golden test sets.
- Task design, checkpoints and verificationVideo lecture, 8′8 min
- Retries, recovery and observabilityVideo lecture, 8′7 min
- Evaluating browser agents: test sets, metrics and benchmarksVideo lecture, 8′7 min
Isolation and least privilege, defending against prompt injection on web pages, and safe handling of credentials, payments and human approval gates.
- Sandboxes, permissions and least privilegeVideo lecture, 8′7 min
- Prompt injection on the webVideo lecture, 8′7 min
- Credentials, payments and human approval gatesVideo lecture, 8′7 min
How MCP, A2A, ACP, UCP and AP2 fit together for agent-to-agent work and agentic commerce, and how to make your own website accurate and usable for AI agents.
- Agent protocols and agentic commerce: MCP, A2A, ACP, UCP and AP2Video lecture, 8′8 min
- Making your website agent-friendlyVideo lecture, 8′7 min
Proven computer-use patterns in marketing operations, QA and data entry, plus the business case, governance and compliance needed to scale from pilot to program.
- Use cases: marketing operations, QA and data entryVideo lecture, 8′7 min
- Business case, governance and compliance for agent programsVideo lecture, 8′7 min
Design, evaluate and govern a browser agent that audits a website with evidence and human approval gates.
- Capstone: a safe browser agent that audits a website with approval gatesVideo lecture, 8′8 min
- Final assessment
Your certificate
Finish with a credential anyone can check
Earn the Certified Computer-Use Agent Engineer badge: The holder can design, build and govern AI agents that operate software through screens and browsers. They understand perception and the agent loop, combine scripts with model reasoning, engineer reliability with verification and evaluation, sandbox agents with least privilege, defend against prompt injection, keep credentials and payments under human approval, and make websites ready for AI agents.
Completed all lessons and scored at least 80% on the final assessment.
- A public verification page
- A PDF certificate to download
- An Open Badge you can share
- One click to your LinkedIn profile
Final assessment
- 25questions drawn from a larger pool
- 40 mintime limit
- 80%pass mark
- 3attempts per 24 hours
Start learning today. It’s free.
Every lesson is free to read. A free account saves your progress, unlocks the final assessment and issues your certificate.