Multimodal & Reasoning Models in PracticeComputer-use and browser agents · Lesson 14 of 17

How computer-use and browser agents work

Article · 11 min · 8 min lecture

Video lecture

How computer-use and browser agents work

11 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 11

Computer-use and browser agents

  • What they are
  • The observe-think-act loop
  • Where they help and struggle
  • Choosing the automation layer

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What they are

Computer-use agents (sometimes called browser agents or GUI agents) operate software through its user interface, the way a person does: they look at the screen, decide what to do, and act by clicking, typing and scrolling. Several AI providers and platforms offer such capabilities, and browser automation products increasingly build them in.

They matter because much business work happens in interfaces without convenient APIs: legacy systems, supplier portals, government websites and internal tools.

The loop

1. Observe: take a screenshot (and sometimes read the page structure or accessibility tree)
2. Think: decide the next action toward the goal
3. Act: click at coordinates or on an element, type text, press keys, scroll, navigate
4. Observe the result, and repeat until done or stopped

This is the agent loop from our agents course, with the screen as the environment. Some systems combine vision with the browser's page structure (the document object model or accessibility tree), which can make element identification more reliable than pure screenshots.

Where they help

  • Filling repetitive forms across portals that lack APIs.
  • Collecting information from multiple websites into a structured summary.
  • Testing web applications by following user journeys.
  • Operating legacy desktop software for data entry or extraction.
  • Assisting users by completing multi-step tasks in a browser on their behalf, with oversight.

Why they are harder than API-based tools

  • Brittleness: layouts change, pop-ups appear, pages load slowly, elements move.
  • Speed and cost: each step involves a screenshot and model call; tasks can take many steps.
  • Ambiguity: similar-looking buttons, hidden menus, dynamic content.
  • Error compounding: one mis-click can derail the rest of the task.
  • Verification: confirming success visually can be unreliable ("Did the form actually submit?").

Capabilities and reliability are improving rapidly, and published benchmark results for these agents change frequently. For any serious deployment, measure on your own tasks.

When to prefer other approaches

Use an API or structured integration (including MCP servers) when one exists: it is faster, cheaper, more reliable and easier to secure. Use traditional scripted automation (such as robotic process automation or browser testing frameworks) when the steps are fixed and the interface is stable. Computer-use agents fit best when the interface varies, the task needs judgement, or no integration exists.

Worked example: supplier portal updates

An e-commerce operations team must update stock availability on five supplier portals weekly, none with an API. A browser agent is given:

  • A goal per portal and a structured list of updates.
  • Login handled by a secure credential mechanism, not typed from the prompt.
  • An instruction to stop before the final submit, take a screenshot of the summary page, and wait for approval.

A team member reviews five summary screenshots and approves. Time saved is substantial, and errors are caught at the review step. After several weeks of clean runs on two stable portals, the team allows those to submit automatically, with daily screenshot audits.

Writing good task instructions for a computer-use agent

Instructions for these agents benefit from being more explicit than for chat tasks, because the agent cannot ask you mid-task:

Goal: update stock levels on the supplier portal for the 12 SKUs listed below.
Start: the portal dashboard (already logged in).
Steps you may need: search each SKU, open its edit page, set "Available qty".
Rules:
- Only change the "Available qty" field. Do not edit prices or descriptions.
- If a SKU is not found, record it and continue.
- If a pop-up asks you to accept new terms, stop and report.
Finish: open the "Pending changes" page, take a screenshot, and stop
without pressing "Submit".
Report: a table of SKU, old qty, new qty, status.

Notice the explicit scope limits, the handling of expected obstacles, the stopping point and the structured report. These turn a vague errand into a checkable job.

Measuring whether it is worth it

Track time per run (agent plus human review) against the manual baseline, error rate caught at review, and how often runs need intervention. If review takes nearly as long as doing the task manually, the design needs rethinking, perhaps by narrowing the task or finding an integration.

Computer use in 2026

Anthropic, OpenAI and Google all offer computer-use or browser-control capabilities through their APIs, and several consumer assistants include agent modes that browse and act on the user's behalf. Tool versions and supported actions change frequently (new tool types, higher-resolution screenshots, zoom actions), so pin the tool version you tested and re-test when upgrading.

Increasingly, agents combine approaches: they call APIs or MCP tools when available, use a browser automation layer (DOM or accessibility tree) for web apps, and fall back to pixel-level computer use only for interfaces with no better option.

Hands-on: the shape of a computer-use loop

The exact tool definitions differ by provider, but every implementation has this shape:

def run_computer_task(goal, model_call, execute_action, take_screenshot,
                      needs_approval, max_steps=50):
    history = [{"role": "user", "content": goal}]
    for step in range(max_steps):
        reply = model_call(history, screenshot=take_screenshot())
        history.append(reply.as_message())
        action = reply.next_action()          # e.g. click(x, y), type(text), key("Enter"), done
        if action is None or action.kind == "done":
            return {"status": "done", "steps": step + 1, "report": reply.text}
        if needs_approval(action):
            return {"status": "paused_for_approval", "action": action.describe()}
        result = execute_action(action)       # runs inside an isolated VM or browser
        history.append(result.as_message())   # includes a fresh screenshot or error
    return {"status": "stopped", "reason": "step limit"}

Everything important lives outside the model: the isolated environment, the approval function, the step limit and the logging. Provider quickstarts (for example Anthropic's computer-use reference implementation, which runs in a container) are the fastest way to try this safely.

Choosing the right automation layer

SituationBest layer
The system has an API or MCP serverAPI or MCP tool calls
Stable web app, fixed stepsScripted browser automation (e.g. Playwright)
Web app that changes, needs judgementBrowser agent using DOM or accessibility tree plus screenshots
Desktop or legacy app, no other accessPixel-level computer use in a VM

Going further

When building or evaluating such agents, design tasks with verifiable end states: a confirmation number captured, a record visible in the system of record, a downloaded file with expected contents. Visual "looks done" is not proof. Log every action with screenshots for audit and debugging.

Key takeaways

  • Computer-use agents observe screens and act via clicks and typing in a loop.
  • They help where no API exists: portals, legacy systems, cross-site research and UI testing.
  • They are slower, costlier and more brittle than API integrations; prefer APIs or scripted automation when available.
  • Design verifiable end states and approval points; log actions with screenshots.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. When is a computer-use agent the best choice?
  2. How should a browser agent confirm a form was successfully submitted?
  3. Which is NOT a typical difficulty for computer-use agents?

Put it into practice

Identify one repetitive browser task in your work. Write the goal, the verifiable end state, the approval point, and whether an API or scripted alternative exists.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.