Computer-Use and Browser Agents: AI That Operates SoftwareHow computer-use agents work · Lesson 1 of 16

From RPA to computer-use agents: what changed

Article · 8 min · 10 min lecture

Video lecture

From RPA to computer-use agents: what changed

15 chapters · about 10 min · full transcript

Coming soon

Chapter 1 of 15

From RPA to computer-use agents

  • What a computer-use agent is
  • The perceive, reason, act loop
  • Where agents beat scripts, and where they don't

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why this matters now

For twenty years, automating software meant one of two things: call an API, or write a brittle script that clicks exact screen positions and CSS selectors. Robotic process automation (RPA) tools made the second option friendlier, but the bots still broke whenever a button moved or a pop-up appeared. Computer-use agents change the economics. A multimodal model looks at the screen (or a structured description of it), decides what to do next, and emits an action such as "click at (512, 742)" or "type this text". Your code executes the action, captures the new state, and the loop repeats until the task is done or a stop condition fires.

The result is automation that can handle software it has never seen before, recover from small surprises and follow instructions written in plain language. It is also slower, more expensive per step and less predictable than a script. This course is about getting the upside without being hurt by the downside.

The perceive, reason, act loop

Every GUI agent, whatever the vendor, runs the same loop:

  1. Perceive. Capture the current state: a screenshot, the page's DOM, the accessibility tree, or a mix.
  2. Reason. The model reads the goal, the history of steps and the current state, then chooses the next action.
  3. Act. Your harness executes the action on a real or virtual machine: mouse, keyboard, navigation, scrolling.
  4. Check. The harness (and ideally a separate verifier) checks whether the action worked and whether the goal is reached.

The model never touches your machine directly. It only proposes actions; your code decides whether to execute them. That single design fact is the foundation of every safety control in this course.

RPA, scripted automation and agents compared

ApproachHow it finds thingsStrengthWeakness
API integrationStructured endpointsFast, reliable, auditableOnly exists where the vendor built one
Scripted browser automation (Playwright, Selenium)CSS/XPath/role selectorsDeterministic, cheap to runBreaks when the UI changes; needs a developer
Classic RPARecorded selectors and coordinatesBusiness-user friendlyBrittle, costly to maintain at scale
Computer-use agentVision plus reasoning over screenshots or accessibility dataAdapts to new and changing UIs; plain-language tasksSlower, costs tokens per step, non-deterministic, new security risks

The practical rule: use the most deterministic tool that can do the job. Prefer an API. If there is no API, prefer a script. Use an agent for the long tail: many different sites, frequently changing UIs, judgment calls, or exploratory work. The best production systems are hybrids, and a later lesson shows how to combine Playwright scripts with model reasoning.

What agents are good at today

Current agents handle tasks such as:

  • Filling multi-step forms from a spreadsheet, where each site's form differs.
  • Checking dozens of landing pages for broken links, missing tracking tags, wrong prices or outdated promotions.
  • Walking through a signup or checkout flow as a QA tester would and reporting what was confusing.
  • Collecting information from sites that have no export (supplier portals, government registries, legacy admin panels).
  • Operating desktop software inside a virtual machine, such as a legacy accounting package with no API.

They still struggle with long tasks where one early mistake poisons everything after it, with dense or unusual interfaces (canvas-heavy apps, complex drag-and-drop), with CAPTCHAs and bot defenses (which they should not try to defeat), and with anything requiring precise timing.

Worked example: a Dubai real-estate agency

An agency in Dubai lists properties on its own site and on several third-party portals. Each portal has a different admin UI and none offers a usable API for the agency's plan. Every Monday a coordinator spends hours checking that prices, photos and availability match the master spreadsheet.

A scripted approach would need a separate brittle script per portal. A computer-use agent receives one instruction ("For each listing ID in this sheet, open the portal listing, compare price, bedrooms and status to the sheet, and report mismatches; do not edit anything"). It runs in a sandboxed browser, logged in with a read-only account where the portal supports one. Output: a mismatch report the coordinator reviews and fixes. Notice the design choices: read-only first, a structured output, and a human who acts on the findings. Only once the report is trusted does the agency consider letting the agent make corrections, and then only behind an approval step.

Where the value really comes from

Teams that succeed with agents rarely start with "let the AI do my job". They start with a narrow, repetitive, verifiable task, measure the baseline (minutes per run, error rate), and build a harness that makes the agent observable. The agent's intelligence matters, but the harness (sandbox, permissions, logging, verification and approval gates) is what makes it deployable.

Pitfalls

  • Automating the wrong layer. If a clean API exists, an agent clicking through the UI is slower and riskier.
  • Unbounded tasks. "Manage our social accounts" is not a task. "Check that the bio link on these five profiles resolves to a 200 page" is.
  • No stop conditions. Agents need step limits, time limits and cost limits.
  • Trusting self-reports. An agent saying "done" is a claim, not evidence. Verify with a screenshot, a database query or a second check.

How to measure success

Track task success rate on a fixed test set, average steps and cost per task, human intervention rate, and time saved versus the manual baseline. You will build this evaluation discipline in module three.

Key takeaways

  • Computer-use agents run a perceive, reason, act, check loop; the model proposes actions and your harness executes them.
  • Prefer APIs, then scripts, then agents; agents shine on the long tail of varied or changing interfaces.
  • Start read-only with structured outputs and a human acting on findings before allowing writes.
  • The harness (sandbox, permissions, logs, verification, approvals) is what makes an agent deployable.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A supplier offers a documented REST API for stock levels. What is the best way to automate stock checks?
  2. Which statement about the model in a computer-use agent is accurate?
  3. Which is the best first task for a new browser agent?

Put it into practice

List three repetitive browser tasks in your team. For each, note whether an API exists, how often the UI changes, and whether the task can start read-only. Pick the best agent candidate.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.