---
title: "How agents see: pixels, the DOM and accessibility trees"
description: "Three ways to perceive a screen An agent can only act as well as it perceives. There are three main representations of what is on screen, and serious…"
url: https://optimizeall.com/learn/computer-use-and-browser-agents/how-agents-see-pixels-dom-accessibility
updated: 2026-10-05
---

Computer-Use and Browser Agents: AI That Operates Software · How computer-use agents work · lesson 2 of 16 · 8 min

# How agents see: pixels, the DOM and accessibility trees

## Three ways to perceive a screen

An agent can only act as well as it perceives. There are three main representations of what is on screen, and serious systems combine them.

**1. Pixels (screenshots).** The harness captures an image and sends it to a vision-capable model. The model identifies buttons, fields and text visually and returns coordinates. This is the most general approach: it works on any desktop app, remote desktop, canvas-rendered chart or legacy system. Anthropic's computer use tool, OpenAI's computer-using agent model and Google's Gemini computer use capability are all screenshot-driven at their core.

**2. The DOM (Document Object Model).** In a browser, the page is a tree of HTML elements. An agent can read the DOM (usually simplified) and target elements by selector or by an index the harness assigns. DOM access gives exact text, link targets and hidden attributes, but raw DOMs are huge and noisy, and they include content that is not visible to the user.

**3. The accessibility tree.** Browsers and operating systems build a simplified tree for screen readers: each node has a **role** (button, link, textbox, heading), an **accessible name** ("Add to cart") and a **state** (checked, expanded, disabled). This is compact, semantic and closer to what a human perceives. Tools such as Playwright's accessibility snapshots, and the Playwright MCP server that exposes them to models, use this representation so a model can say "click the button named Submit" instead of guessing pixels.

## Trade-offs at a glance

| Representation | Works on | Precision | Token cost | Typical failure |
|---|---|---|---|---|
| Screenshot | Anything with a screen | Coordinates can be off by a few pixels | Image tokens per step | Misclicks on small or crowded targets |
| DOM | Web pages | Exact elements and attributes | Can be very large | Noise, invisible elements, shadow DOM, iframes |
| Accessibility tree | Web and native apps with good accessibility | Semantic roles and names | Compact | Poorly labeled sites expose unnamed "button" nodes |

A key insight for marketers and site owners: **the accessibility tree is the agent's view of your website.** If your "Book a demo" control is an unlabeled div with a click handler, screen readers struggle and so do agents. Good accessibility (WCAG) and agent-friendliness overlap heavily, which you will use in the module on agent-friendly websites.

## Coordinates and scaling

Screenshot-based agents return coordinates in the pixel space of the image they were shown. If your harness downscales a 2560×1440 screenshot to fit the model's recommended image size, it must scale the model's coordinates back up before clicking. Getting this wrong produces agents that consistently click slightly beside the target. Vendors publish recommended maximum image sizes; check the current documentation for your model, because these limits change between model generations. Some tools use normalized coordinates on a fixed grid that you convert to real pixels, so read the docs for the tool you use.

Useful tactics:

- Keep the virtual display at a moderate, fixed resolution so screenshots need little or no scaling.
- Use a **zoom** action where the tool provides one, to inspect a small region at full resolution before clicking tiny controls.
- Prefer keyboard actions (Tab, Enter, shortcuts) when targets are small.

## Hybrid perception in practice

Many production browser agents send both a screenshot and a trimmed accessibility snapshot. The screenshot provides layout and visual cues (which modal is on top, what is grayed out); the tree provides exact names to target. When the tree is poor, the agent falls back to vision.

## Hands-on: see what an agent sees

Install Playwright for Python and print a page's accessibility snapshot next to a screenshot. This is the fastest way to understand why some sites are easy for agents and others are not.

```bash
python -m venv .venv && source .venv/bin/activate
pip install playwright
playwright install chromium
```

```python
# see_like_an_agent.py
import sys
from playwright.sync_api import sync_playwright

url = sys.argv[1] if len(sys.argv) > 1 else "https://example.com"

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(viewport={"width": 1280, "height": 800})
    try:
        page.goto(url, wait_until="domcontentloaded", timeout=30000)
        page.screenshot(path="view.png")
        # ARIA snapshot: a YAML-like outline of roles and accessible names
        print(page.locator("body").aria_snapshot())
    except Exception as exc:
        print(f"Failed to load {url}: {exc}")
    finally:
        browser.close()
```

Run it against your own homepage. Look for nodes like `button` with no name, links whose name is just "click here", or form fields without labels. Each of those is a place where an agent (and a screen-reader user) must guess.

## Worked example: a Lahore e-commerce checkout

A Lahore fashion retailer tested a browser agent on its checkout. It kept failing to select a delivery city. The screenshot showed a styled dropdown; the accessibility tree showed a generic element with no role or name, because the dropdown was a custom component. The agent clicked, the list rendered in a portal at the bottom of the DOM, and the agent lost track. Replacing it with an accessible combobox (proper role, label and keyboard support) fixed the agent and also improved mobile and screen-reader usability.

## Pitfalls

- Treating DOM text as visible truth: hidden elements and off-screen text can mislead the agent, and attackers can hide instructions there (module four).
- Sending the full raw DOM to the model: expensive and noisy. Trim to interactive and visible elements.
- Forgetting iframes and shadow DOM, which many payment and chat widgets use.

## How to measure success

For a set of target pages, measure how often the agent identifies the correct element on the first attempt, the misclick rate, and tokens per step for each perception mode. Choose the cheapest mode that meets your accuracy bar.

## Video lecture: How agents see: pixels, the DOM and accessibility trees

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. How agents see
2. 1. Pixels
3. Why it matters
4. Describing a room by phone
5. Simple example: sign-up form
6. 2. DOM and 3. Accessibility tree
7. Your site through an agent's eyes
8. Coordinate scaling
9. Worked example: a checkout dropdown
10. Hybrid perception
11. Quick self-check
12. Iframes and shadow DOM
13. Three mistakes
14. Try this now
15. Recap

## Lecture transcript

### How agents see

An agent can only act as well as it sees. And here is something most people never think about: when an AI agent visits your website, it probably does not see it the way you do. In this lesson you will learn the three ways agents perceive a screen, why each one fails in its own way, and how to look at your own site through an agent's eyes in about five minutes.

### 1. Pixels

The first way is pixels. Your harness takes a screenshot and sends it to a vision-capable model. The model spots buttons and fields visually and replies with coordinates to click. This is the most general approach. It works on desktop apps, remote desktops, charts drawn on a canvas, even ancient accounting software. The big computer-use tools from Anthropic, OpenAI and Google are screenshot-driven at their core. The weakness is precision. Small or crowded targets get misclicked.

### Why it matters

Why does this matter? Because perception is where most agent failures begin. If the agent misreads the screen, every decision after that is built on sand. And it matters for site owners too: as more people send AI assistants to compare products, fill forms and book services, the way your site appears to an agent affects whether those tasks succeed. Understanding perception helps you build better agents and better websites at the same time.

### Describing a room by phone

Here's an analogy. Imagine describing a room to a friend over the phone. You could send a photo, that's the screenshot: rich, but they have to work out where everything is. You could read out the architect's full blueprint, that's the DOM: precise, but overwhelming, and it includes pipes hidden behind the walls. Or you could list the furniture with labels, sofa, lamp, door to the kitchen, that's the accessibility tree: short, meaningful and easy to act on. Good agents, like good friends on the phone, often use the photo and the labeled list together.

### Simple example: sign-up form

Let's walk through a simple example with a newsletter sign-up form. On screen, you see an email box and a bright button. In a good accessibility tree, you'll see: textbox, named Email address, and button, named Subscribe. An agent can simply say fill the textbox named Email address and click Subscribe. Now picture the same form built with a placeholder instead of a label and an icon-only button. The tree shows textbox with no name, and button with no name. The agent has to guess from pixels, and guesses are where misclicks come from.

### 2. DOM and 3. Accessibility tree

The second way is the DOM, the tree of HTML elements behind every web page. Reading the DOM gives exact text, link targets and attributes. But raw DOMs are enormous and noisy, and they contain things users never see, such as hidden menus and off-screen text. The third way is the accessibility tree. Browsers build it for screen readers. Every node has a role, like button or textbox, a name, like Add to cart, and a state, like disabled. It is compact and meaningful, which is why tools such as Playwright's accessibility snapshots, and the Playwright MCP server, hand it to models.

### Your site through an agent's eyes

Here is the idea to take away for your own business. The accessibility tree is the agent's view of your website. If your Book a demo button is really an unlabeled box with a click handler, screen-reader users struggle, and so do agents. Good accessibility and agent-friendliness overlap heavily. Fixing labels, roles and keyboard support helps disabled users, helps agents, and usually helps your conversion rate too.

### Coordinate scaling

Now a practical trap. Screenshot agents return coordinates in the pixel space of the image they saw. If your harness shrinks a large screenshot to fit the model's recommended size, and forgets to scale the coordinates back up, every click lands slightly off. The fixes are simple. Keep the virtual display at a moderate fixed resolution. Use a zoom action, where your tool offers one, before clicking tiny controls. And prefer the keyboard, Tab and Enter, when targets are small. Recommended image sizes change between model generations, so check the current docs.

### Worked example: a checkout dropdown

A story from Lahore. A fashion retailer tested an agent on its checkout, and it kept failing to pick a delivery city. The screenshot showed a pretty dropdown. The accessibility tree showed a nameless generic element, because the dropdown was a custom component that rendered its list somewhere else on the page. The agent clicked, the list appeared far away in the DOM, and the agent lost the thread. The team swapped in a proper accessible combobox. The agent started succeeding, and mobile users and screen-reader users had an easier time as well.

### Hybrid perception

Most production browser agents combine these views. They send a screenshot for layout, such as which pop-up is on top, and a trimmed accessibility snapshot for exact names to target. When the tree is poor, they fall back to vision. Two cautions. Never send the full raw DOM, trim it to visible, interactive elements. And remember that hidden text in the DOM can mislead an agent, or even carry an attacker's instructions, which we will deal with in the security module.

### Quick self-check

Here's a quick exercise you can do in your head right now. Picture your own website's main call to action. Is it a real button or link, with text that says exactly what happens, like Get a quote? Or is it an icon, or a styled box with the words click here? Now picture your contact form. Does every field have a visible label, or do you rely on placeholder text that disappears when someone starts typing? Each of those details decides whether an agent can act confidently or has to guess.

### Iframes and shadow DOM

One more subtlety: iframes and shadow DOM. Many payment forms, chat widgets and embedded booking tools live inside iframes or web components. A naive agent that only reads the top-level page won't see them at all, while a screenshot shows them clearly. That mismatch is a classic source of confusion. Good harnesses handle frames explicitly, and good websites keep critical steps, like choosing a time slot, accessible rather than buried in a third-party widget with no labels.

### Three mistakes

Three mistakes to avoid. First, treating everything in the DOM as visible truth. Hidden elements and off-screen text can mislead the agent, and attackers can hide instructions there, which we'll tackle in the security module. Second, sending the full raw DOM to the model. It's expensive and noisy. Trim it to visible, interactive elements. Third, forgetting iframes and shadow DOM. Payment forms and chat widgets often live there, and a harness that ignores them will swear an element doesn't exist when it's right there on the screen.

### Try this now

Try this now. Install Playwright using the commands in the lesson text and run the short script against your own homepage and your most important form page. Look at the accessibility snapshot it prints. Count three things: buttons with no name, form fields without labels, and links whose text is something vague like click here or read more. Write each one down with the page it's on. Those are the exact places where both agents and screen-reader users are guessing, and fixing them is usually quick.

### Recap

Let's recap. Agents perceive through pixels, the DOM, or the accessibility tree, and good systems blend them. The accessibility tree is effectively how agents read your site. Scale coordinates correctly, and use zoom and keyboard actions for small targets. Your next step: run the short Playwright script in the lesson text on your homepage, and list every unnamed button or unlabeled field it reveals. Next up, we will write the agent loop itself.

## Key takeaways

- Agents perceive via screenshots, the DOM or the accessibility tree; robust systems combine them.
- The accessibility tree (roles, names, states) is effectively the agent's view of your site.
- Screenshot coordinates must be scaled back if you resize images; use zoom and keyboard actions for small targets.
- Trim DOM input to visible, interactive elements and watch for hidden text, iframes and shadow DOM.

## Try it

Run the hands-on script on your homepage and one key conversion page. List every unnamed button, unlabeled field or vague link text in the snapshot.

- [Previous: From RPA to computer-use agents: what changed](https://optimizeall.com/learn/computer-use-and-browser-agents/from-rpa-to-computer-use-agents)
- [Next: Writing the agent loop: a minimal computer-use harness](https://optimizeall.com/learn/computer-use-and-browser-agents/writing-the-agent-loop)
- [All lessons of Computer-Use and Browser Agents: AI That Operates Software](https://optimizeall.com/learn/computer-use-and-browser-agents)
