Computer-Use and Browser Agents: AI That Operates SoftwareHow computer-use agents work · Lesson 2 of 16

How agents see: pixels, the DOM and accessibility trees

Article · 8 min · 9 min lecture

Video lecture

How agents see: pixels, the DOM and accessibility trees

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

How agents see

  • Pixels
  • The DOM
  • The accessibility tree
  • Seeing your site as an agent does

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Three ways to perceive a screen

An agent can only act as well as it perceives. There are three main representations of what is on screen, and serious systems combine them.

1. Pixels (screenshots). The harness captures an image and sends it to a vision-capable model. The model identifies buttons, fields and text visually and returns coordinates. This is the most general approach: it works on any desktop app, remote desktop, canvas-rendered chart or legacy system. Anthropic's computer use tool, OpenAI's computer-using agent model and Google's Gemini computer use capability are all screenshot-driven at their core.

2. The DOM (Document Object Model). In a browser, the page is a tree of HTML elements. An agent can read the DOM (usually simplified) and target elements by selector or by an index the harness assigns. DOM access gives exact text, link targets and hidden attributes, but raw DOMs are huge and noisy, and they include content that is not visible to the user.

3. The accessibility tree. Browsers and operating systems build a simplified tree for screen readers: each node has a role (button, link, textbox, heading), an accessible name ("Add to cart") and a state (checked, expanded, disabled). This is compact, semantic and closer to what a human perceives. Tools such as Playwright's accessibility snapshots, and the Playwright MCP server that exposes them to models, use this representation so a model can say "click the button named Submit" instead of guessing pixels.

Trade-offs at a glance

RepresentationWorks onPrecisionToken costTypical failure
ScreenshotAnything with a screenCoordinates can be off by a few pixelsImage tokens per stepMisclicks on small or crowded targets
DOMWeb pagesExact elements and attributesCan be very largeNoise, invisible elements, shadow DOM, iframes
Accessibility treeWeb and native apps with good accessibilitySemantic roles and namesCompactPoorly labeled sites expose unnamed "button" nodes

A key insight for marketers and site owners: the accessibility tree is the agent's view of your website. If your "Book a demo" control is an unlabeled div with a click handler, screen readers struggle and so do agents. Good accessibility (WCAG) and agent-friendliness overlap heavily, which you will use in the module on agent-friendly websites.

Coordinates and scaling

Screenshot-based agents return coordinates in the pixel space of the image they were shown. If your harness downscales a 2560×1440 screenshot to fit the model's recommended image size, it must scale the model's coordinates back up before clicking. Getting this wrong produces agents that consistently click slightly beside the target. Vendors publish recommended maximum image sizes; check the current documentation for your model, because these limits change between model generations. Some tools use normalized coordinates on a fixed grid that you convert to real pixels, so read the docs for the tool you use.

Useful tactics:

  • Keep the virtual display at a moderate, fixed resolution so screenshots need little or no scaling.
  • Use a zoom action where the tool provides one, to inspect a small region at full resolution before clicking tiny controls.
  • Prefer keyboard actions (Tab, Enter, shortcuts) when targets are small.

Hybrid perception in practice

Many production browser agents send both a screenshot and a trimmed accessibility snapshot. The screenshot provides layout and visual cues (which modal is on top, what is grayed out); the tree provides exact names to target. When the tree is poor, the agent falls back to vision.

Hands-on: see what an agent sees

Install Playwright for Python and print a page's accessibility snapshot next to a screenshot. This is the fastest way to understand why some sites are easy for agents and others are not.

python -m venv .venv && source .venv/bin/activate
pip install playwright
playwright install chromium
# see_like_an_agent.py
import sys
from playwright.sync_api import sync_playwright

url = sys.argv[1] if len(sys.argv) > 1 else "https://example.com"

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(viewport={"width": 1280, "height": 800})
    try:
        page.goto(url, wait_until="domcontentloaded", timeout=30000)
        page.screenshot(path="view.png")
        # ARIA snapshot: a YAML-like outline of roles and accessible names
        print(page.locator("body").aria_snapshot())
    except Exception as exc:
        print(f"Failed to load {url}: {exc}")
    finally:
        browser.close()

Run it against your own homepage. Look for nodes like button with no name, links whose name is just "click here", or form fields without labels. Each of those is a place where an agent (and a screen-reader user) must guess.

Worked example: a Lahore e-commerce checkout

A Lahore fashion retailer tested a browser agent on its checkout. It kept failing to select a delivery city. The screenshot showed a styled dropdown; the accessibility tree showed a generic element with no role or name, because the dropdown was a custom component. The agent clicked, the list rendered in a portal at the bottom of the DOM, and the agent lost track. Replacing it with an accessible combobox (proper role, label and keyboard support) fixed the agent and also improved mobile and screen-reader usability.

Pitfalls

  • Treating DOM text as visible truth: hidden elements and off-screen text can mislead the agent, and attackers can hide instructions there (module four).
  • Sending the full raw DOM to the model: expensive and noisy. Trim to interactive and visible elements.
  • Forgetting iframes and shadow DOM, which many payment and chat widgets use.

How to measure success

For a set of target pages, measure how often the agent identifies the correct element on the first attempt, the misclick rate, and tokens per step for each perception mode. Choose the cheapest mode that meets your accuracy bar.

Key takeaways

  • Agents perceive via screenshots, the DOM or the accessibility tree; robust systems combine them.
  • The accessibility tree (roles, names, states) is effectively the agent's view of your site.
  • Screenshot coordinates must be scaled back if you resize images; use zoom and keyboard actions for small targets.
  • Trim DOM input to visible, interactive elements and watch for hidden text, iframes and shadow DOM.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An agent clicks consistently about 20% to the left and above the intended buttons. What is the most likely cause?
  2. Which representation gives compact, semantic information such as 'button named Add to cart'?
  3. Why does improving WCAG accessibility often make a site easier for agents?

Put it into practice

Run the hands-on script on your homepage and one key conversion page. List every unnamed button, unlabeled field or vague link text in the snapshot.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.