Computer-Use and Browser Agents: AI That Operates SoftwarePlatforms and hybrid automation · Lesson 4 of 16
The computer-use landscape: APIs, agent products and AI browsers
Video lecture
The computer-use landscape: APIs, agent products and AI browsers
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 The computer-use landscape
Here is a problem with learning about computer-use agents. By the time you finish a course, half the product names have changed. Between early 2025 and late 2026, the big vendors launched, renamed, merged and retired agent products several times. So in this lesson, you will learn the landscape as categories, and you will leave with a checklist that still works when the next product launches.
0:29 1. Model APIs
Category one is model APIs with a computer-use tool. You provide the environment, a virtual machine, a container or a browser, and the model proposes actions. Anthropic's Claude API has a generally available computer toolset for recent models, plus a reference Docker demo. OpenAI's computer-using agent model, which powered Operator, is available as a computer use tool in its Responses API, with safety checks you must acknowledge. Google offers computer use through the Gemini API, starting with a Gemini two point five Computer Use preview in October 2025. This is the category you use when a business process must be controlled and audited.
1:14 Why it matters
Why does this matter? Because the choice of tool decides where your data goes, what the agent can reach, and whether you can prove what it did. Pick a consumer product for a regulated process and you may not be able to log or restrict it properly. Pick a heavy custom build for an ad-hoc research task and you'll waste weeks. Matching the category to the job saves money, protects data and avoids painful migrations when the market shifts again.
1:49 Like early smartphones
Here's an analogy for this market. Think of smartphones around the late two thousands. New models, new names and new features appeared every few months, and many disappeared just as fast. The people who navigated that well didn't memorize model numbers. They knew what to look for: battery, camera, security updates, app support. Computer-use tools are at the same stage. Model names and product brands will keep changing. Your evaluation questions shouldn't.
2:20 Simple example: an AI browser trial
A simple example of using the checklist. Say a colleague wants to try an AI browser to research competitors' pricing pages. Run the first three questions. Where does it run? In their everyday browser. What can it reach? Everything they're logged into, including the company's ad accounts. Can we restrict domains in code? Not really. So the answer is: fine to try, but in a separate browser profile with no work logins, and nothing sensitive. Three questions, one minute, and a much safer experiment.
2:57 2. Agent products
Category two is finished agent products, where the vendor runs a cloud browser for you. The OpenAI story shows how fluid this is. Operator launched as a research preview in January 2025, was folded into ChatGPT as agent mode that July, and has been reorganized since. Google's Project Mariner prototype fed into Gemini's agent features and AI Mode in Search. These are great for individuals and one-off tasks. For repeatable business processes, you usually want your own harness.
3:31 3. AI browsers
Category three is AI browsers and browser assistants. Think Perplexity's Comet, Gemini features in Chrome including auto browse on paid plans in supported regions, AI built into Microsoft Edge, and extensions like Claude in Chrome. These work inside your real, logged-in browser. That is powerful, because the agent can use your accounts. It is also risky, for exactly the same reason. A malicious page can try to hijack an agent that has access to your email and admin panels.
4:05 4. Open source and infrastructure
Category four is open source and infrastructure. Playwright remains the standard for scripted browser automation, and the Playwright MCP server lets any MCP-capable model drive a browser using accessibility snapshots. Libraries such as Browser Use and Stagehand put model reasoning on top of browser automation. And hosted browser providers run fleets of isolated cloud browsers with session recording, so you do not have to operate your own.
4:34 Evaluation checklist
Now the checklist. Where does the environment run? What can the agent reach? Can you restrict actions and domains in code, not just in a prompt? Is there a built-in human confirmation step? What defenses exist against prompt injection? Can you log and replay every step? What are the data retention and training terms, which matters under GDPR, UK GDPR, Saudi Arabia's PDPL and the UAE's data law? What does it cost per task on your workload? And is it generally available, or still a preview?
5:11 Worked example: one agency, two choices
Here is how a Manchester agency used it. Account managers wanted ad-hoc research help across client sites. For that, an AI browser was fine, but in a separate profile with no client admin logins. They also needed a weekly audit of two hundred client landing pages. For that, they built an API-based harness with a headless cloud browser, a domain allowlist per client, read-only credentials and stored evidence, because clients want to see exactly what was checked and when.
5:45 Run a fair pilot
How do you run a fair pilot between two tools? Take twenty real tasks from your workflow. Run each on both tools, ideally more than once. For every run, record whether it succeeded when checked, how many times a person had to step in, how long it took, and what it cost. Keep the environment the same for both. At the end you'll have a simple table, and the decision usually becomes obvious. It almost never matches the impression you got from the demo videos.
6:22 Data terms matter
A word on data terms, because this is where legal and IT teams will stop you. When an agent sends screenshots to a model provider, those screenshots may contain customer names, emails or financial details. Check how long the provider keeps that data, whether it's used for training, where it's processed, and whether enterprise or zero-retention options exist. For teams in the UK, the EU, Saudi Arabia or the UAE, cross-border processing questions come up quickly. Settle them before the pilot, not after.
6:58 Three mistakes
Three mistakes in choosing tools. First, choosing by demo video, which is designed to succeed. Always pilot on your own tasks. Second, building a business process on a preview feature with no fallback. If the vendor changes or withdraws it, you need a plan. Third, ignoring data terms until legal or IT stops the project. Screenshots can contain personal data, so settle retention, training use and processing region at the start.
7:29 Try this now
Try this now. Take one real task from your work and fill in the evaluation checklist for two options: one API-based build and one finished product or AI browser. For each checklist question, write a one-line answer and where you found it, ideally in the vendor's documentation. Where you can't find an answer, write unknown. Then run each option on the same five examples of the task, and record success, human interventions, time and cost. You'll have a small, honest comparison instead of an opinion.
8:06 Recap
To recap. Learn categories, not names: model APIs, agent products, AI browsers, and open-source infrastructure. Use products for ad-hoc work and your own harness for repeatable processes. And evaluate every tool with the same checklist. Your next step: fill in the checklist for two tools you could use this week, against one real task. Next, we combine deterministic Playwright scripts with model reasoning to get the best of both.
Why the names keep changing
Few areas of AI move faster than computer use. Between January 2025 and September 2026, major vendors launched, renamed, merged and retired agent products several times. If you learn product names, your knowledge expires in months. If you learn categories and evaluation criteria, it lasts. This lesson maps the landscape as of September 2026 and gives you a checklist for evaluating whatever launches next. Treat every specific here as "verify in the current docs before you build".
Category 1: Model APIs with a computer-use tool
These let developers build their own agents. You supply the environment (a VM, container or browser), the model proposes actions.
- Anthropic (Claude API). Computer use was first released as a public beta in October 2024. As of September 2026 the Claude API offers a generally available computer toolset for recent Claude models, with dated beta versions still used on some cloud platforms such as Amazon Bedrock. Anthropic also publishes a Dockerized reference implementation. Related tools (bash, a text editor) let the same agent work in terminals and files.
- OpenAI (Responses API). OpenAI's computer-using agent (CUA) model powered Operator, a research preview launched in January 2025, and is exposed to developers as a computer use tool in the Responses API, with an official sample app. The API includes safety checks your code must acknowledge.
- Google (Gemini API). Google released the Gemini 2.5 Computer Use model in public preview in October 2025, via Google AI Studio and Vertex AI, optimized for browsers with some mobile support. Newer Gemini models expose computer use as a tool. Google states that preview capabilities may be error-prone and recommends close supervision for important tasks.
When you compare these, look at supported environments (browser only vs full desktop), action set, coordinate model, safety features (built-in injection classifiers, confirmation prompts), latency, cost per step, and data-handling terms.
Category 2: Agent products for end users
These are finished products where the vendor runs the environment, usually a cloud browser. The OpenAI lineage shows how fluid this is: Operator (January 2025) was folded into ChatGPT as "agent mode" in July 2025, and OpenAI has continued to reorganize its agentic and browser features since. Google's research prototype Project Mariner fed its technology into Gemini's agent features and AI Mode in Search, and Google has since folded the prototype into its Gemini products. Anthropic ships Claude in Chrome, a browser extension that lets Claude act in the user's own browser, and desktop computer-use features in its apps.
The durable lesson: consumer agent products are convenient for individuals and ad-hoc tasks, but for repeatable business processes you usually want an API-based harness you control, with your own logging, permissions and evaluation.
Category 3: AI browsers and browser assistants
A separate wave puts the agent inside the browser itself. Examples as of 2026 include Perplexity's Comet browser, Gemini features in Chrome (including an "auto browse" capability for paid Google AI plans in supported regions), AI features built into Microsoft Edge, and extensions such as Claude in Chrome. These act inside the user's real, logged-in session, which is powerful and risky: the agent can reach your email, bank and admin panels, and malicious pages can try to hijack it.
Category 4: Open-source frameworks and infrastructure
- Playwright (Microsoft) remains the standard for scripted browser automation, and the Playwright MCP server exposes browser control to any MCP-capable model using accessibility snapshots.
- Open-source agent libraries such as Browser Use and Stagehand add LLM reasoning on top of browser automation, with different balances between deterministic code and model decisions.
- Hosted browser infrastructure providers run fleets of isolated cloud browsers with session recording, which saves you operating your own.
An evaluation checklist that outlives product names
| Question | Why it matters |
|---|---|
| Where does the environment run: your VM, vendor cloud, or user's own browser? | Determines data exposure and blast radius |
| What can the agent reach: which sites, files, accounts? | Least privilege |
| Can you restrict actions and domains in code? | Real control vs prompt-only control |
| Is there a built-in human confirmation mechanism? | Needed for purchases, sends, deletes |
| What defenses exist against prompt injection? | Web pages are untrusted input |
| Can you log, replay and export every step? | Audit, debugging, compliance |
| What are the data retention and training terms? | Privacy obligations (GDPR, UK GDPR, PDPL in KSA, UAE PDPL) |
| What is the cost and latency per task on your workload? | ROI; measure, don't assume |
| Is it GA or preview? | Preview features may change or disappear |
Worked example: choosing for a UK marketing agency
A Manchester agency wants two things: (a) account managers doing ad-hoc research across client sites, and (b) a weekly automated audit of 200 client landing pages. For (a), a consumer agent product or AI browser is fine in a separate browser profile with no client admin logins. For (b), they build an API-based harness with a headless cloud browser, a domain allowlist per client, read-only credentials and stored evidence, because they must show clients exactly what was checked and when.
Pitfalls
- Choosing by demo video. Always run a pilot on your own tasks with a fixed test set.
- Building a process on a preview feature with no fallback.
- Letting staff run AI browsers in the same profile as company admin accounts.
How to measure success
Pick two candidate platforms and run the same 20 tasks on each. Record success rate, human interventions, cost and duration. Decide with data.
Key takeaways
- Learn categories and criteria, not product names; this market renames and retires products frequently.
- Model APIs (Anthropic, OpenAI, Google) let you build controlled harnesses; consumer agents and AI browsers suit ad-hoc work.
- AI browsers act in real logged-in sessions, which increases both power and blast radius.
- Evaluate on environment, reach, code-level restrictions, confirmations, injection defenses, logging, data terms, cost and GA status.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Fill the evaluation checklist for two tools you could use today (one API-based, one product or AI browser) against a task from your own work.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.