Computer-Use and Browser Agents: AI That Operates SoftwarePlatforms and hybrid automation · Lesson 4 of 16

The computer-use landscape: APIs, agent products and AI browsers

Article · 8 min · 8 min lecture

Video lecture

The computer-use landscape: APIs, agent products and AI browsers

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

The computer-use landscape

  • Why names keep changing
  • Four categories
  • A checklist that outlives product names

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why the names keep changing

Few areas of AI move faster than computer use. Between January 2025 and September 2026, major vendors launched, renamed, merged and retired agent products several times. If you learn product names, your knowledge expires in months. If you learn categories and evaluation criteria, it lasts. This lesson maps the landscape as of September 2026 and gives you a checklist for evaluating whatever launches next. Treat every specific here as "verify in the current docs before you build".

Category 1: Model APIs with a computer-use tool

These let developers build their own agents. You supply the environment (a VM, container or browser), the model proposes actions.

  • Anthropic (Claude API). Computer use was first released as a public beta in October 2024. As of September 2026 the Claude API offers a generally available computer toolset for recent Claude models, with dated beta versions still used on some cloud platforms such as Amazon Bedrock. Anthropic also publishes a Dockerized reference implementation. Related tools (bash, a text editor) let the same agent work in terminals and files.
  • OpenAI (Responses API). OpenAI's computer-using agent (CUA) model powered Operator, a research preview launched in January 2025, and is exposed to developers as a computer use tool in the Responses API, with an official sample app. The API includes safety checks your code must acknowledge.
  • Google (Gemini API). Google released the Gemini 2.5 Computer Use model in public preview in October 2025, via Google AI Studio and Vertex AI, optimized for browsers with some mobile support. Newer Gemini models expose computer use as a tool. Google states that preview capabilities may be error-prone and recommends close supervision for important tasks.

When you compare these, look at supported environments (browser only vs full desktop), action set, coordinate model, safety features (built-in injection classifiers, confirmation prompts), latency, cost per step, and data-handling terms.

Category 2: Agent products for end users

These are finished products where the vendor runs the environment, usually a cloud browser. The OpenAI lineage shows how fluid this is: Operator (January 2025) was folded into ChatGPT as "agent mode" in July 2025, and OpenAI has continued to reorganize its agentic and browser features since. Google's research prototype Project Mariner fed its technology into Gemini's agent features and AI Mode in Search, and Google has since folded the prototype into its Gemini products. Anthropic ships Claude in Chrome, a browser extension that lets Claude act in the user's own browser, and desktop computer-use features in its apps.

The durable lesson: consumer agent products are convenient for individuals and ad-hoc tasks, but for repeatable business processes you usually want an API-based harness you control, with your own logging, permissions and evaluation.

Category 3: AI browsers and browser assistants

A separate wave puts the agent inside the browser itself. Examples as of 2026 include Perplexity's Comet browser, Gemini features in Chrome (including an "auto browse" capability for paid Google AI plans in supported regions), AI features built into Microsoft Edge, and extensions such as Claude in Chrome. These act inside the user's real, logged-in session, which is powerful and risky: the agent can reach your email, bank and admin panels, and malicious pages can try to hijack it.

Category 4: Open-source frameworks and infrastructure

  • Playwright (Microsoft) remains the standard for scripted browser automation, and the Playwright MCP server exposes browser control to any MCP-capable model using accessibility snapshots.
  • Open-source agent libraries such as Browser Use and Stagehand add LLM reasoning on top of browser automation, with different balances between deterministic code and model decisions.
  • Hosted browser infrastructure providers run fleets of isolated cloud browsers with session recording, which saves you operating your own.

An evaluation checklist that outlives product names

QuestionWhy it matters
Where does the environment run: your VM, vendor cloud, or user's own browser?Determines data exposure and blast radius
What can the agent reach: which sites, files, accounts?Least privilege
Can you restrict actions and domains in code?Real control vs prompt-only control
Is there a built-in human confirmation mechanism?Needed for purchases, sends, deletes
What defenses exist against prompt injection?Web pages are untrusted input
Can you log, replay and export every step?Audit, debugging, compliance
What are the data retention and training terms?Privacy obligations (GDPR, UK GDPR, PDPL in KSA, UAE PDPL)
What is the cost and latency per task on your workload?ROI; measure, don't assume
Is it GA or preview?Preview features may change or disappear

Worked example: choosing for a UK marketing agency

A Manchester agency wants two things: (a) account managers doing ad-hoc research across client sites, and (b) a weekly automated audit of 200 client landing pages. For (a), a consumer agent product or AI browser is fine in a separate browser profile with no client admin logins. For (b), they build an API-based harness with a headless cloud browser, a domain allowlist per client, read-only credentials and stored evidence, because they must show clients exactly what was checked and when.

Pitfalls

  • Choosing by demo video. Always run a pilot on your own tasks with a fixed test set.
  • Building a process on a preview feature with no fallback.
  • Letting staff run AI browsers in the same profile as company admin accounts.

How to measure success

Pick two candidate platforms and run the same 20 tasks on each. Record success rate, human interventions, cost and duration. Decide with data.

Key takeaways

  • Learn categories and criteria, not product names; this market renames and retires products frequently.
  • Model APIs (Anthropic, OpenAI, Google) let you build controlled harnesses; consumer agents and AI browsers suit ad-hoc work.
  • AI browsers act in real logged-in sessions, which increases both power and blast radius.
  • Evaluate on environment, reach, code-level restrictions, confirmations, injection defenses, logging, data terms, cost and GA status.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An agency needs a weekly, auditable check of 200 client landing pages. What fits best?
  2. Why is running an AI browser agent in your everyday logged-in browser profile risky?
  3. Which question best protects you from building on a moving target?

Put it into practice

Fill the evaluation checklist for two tools you could use today (one API-based, one product or AI browser) against a task from your own work.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.