Computer-Use and Browser Agents: AI That Operates SoftwareSandboxing, permissions and security · Lesson 9 of 16

Sandboxes, permissions and least privilege

Article · 7 min · 8 min lecture

Video lecture

Sandboxes, permissions and least privilege

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Sandboxes and least privilege

  • Assume mistakes will happen
  • Isolation
  • Least privilege
  • A locked-down sandbox

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Assume the agent will make mistakes

Design every computer-use deployment on one assumption: at some point the agent will do something you did not intend, either through its own error or because a web page manipulated it. Security is about making that moment boring. The two tools are isolation (the agent can only touch a contained environment) and least privilege (inside that environment, it can only do what the task requires).

Vendor documentation says the same. Anthropic's computer use guidance recommends a dedicated virtual machine or container with minimal privileges, avoiding access to sensitive data and credentials, limiting internet access to an allowlist of domains, and asking a human to confirm consequential actions. Google's and OpenAI's guidance similarly recommends close supervision and confirmation for sensitive actions.

Layers of isolation

LayerOptionsWhat it contains
ComputeDedicated VM, container, or hosted cloud browserThe agent cannot reach your laptop, files or corporate network
Browser profileFresh, empty profile per run; no saved passwords or synced accountsNo accidental access to your email, banking or admin sessions
NetworkEgress proxy with domain allowlist; block internal IP rangesAgent cannot wander to arbitrary sites or internal services
IdentityDedicated service accounts with minimal rolesMistakes are limited to what that role can do
DataOnly the inputs needed for this taskNothing extra to leak

A container is usually enough for browser automation: a headless Chromium in a container with no host mounts, a non-root user and an egress proxy. For full desktop agents, a VM or disposable cloud desktop provides stronger separation. Destroy the environment after each run, or at least reset it, so nothing persists between tasks or clients.

Hands-on: a locked-down browser container

# Dockerfile: headless browser agent sandbox (illustrative; pin versions you test)
FROM mcr.microsoft.com/playwright/python:latest
RUN useradd -m agent
USER agent
WORKDIR /home/agent/app
COPY --chown=agent:agent requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY --chown=agent:agent . .
ENV HTTPS_PROXY=http://egress-proxy:3128 HTTP_PROXY=http://egress-proxy:3128
CMD ["python", "run_task.py"]
# docker-compose.yml: agent can only reach the internet through the allowlisting proxy
services:
  agent:
    build: .
    env_file: .env            # API keys come from the environment, never baked into the image
    read_only: true
    tmpfs: ["/tmp", "/home/agent/.cache"]
    cap_drop: ["ALL"]
    security_opt: ["no-new-privileges:true"]
    networks: [sandbox]
  egress-proxy:
    image: ubuntu/squid:latest
    volumes: ["./squid.conf:/etc/squid/squid.conf:ro"]
    networks: [sandbox, outside]
networks:
  sandbox: { internal: true }   # no direct internet from the agent container
  outside: {}
# squid.conf (excerpt): allow only the client's domains
acl allowed dstdomain .example-client.com .example-cdn.net
http_access allow allowed
http_access deny all

Pin image versions you have tested rather than latest in production. Also enforce the allowlist in your harness (refuse goto to other domains) so you have two independent controls.

Least privilege for accounts

  • Create dedicated agent accounts on each system with the smallest role that works: viewer or analyst rather than admin; one client, not all clients.
  • Prefer read-only roles; add write permissions only for specific, reviewed workflows.
  • Use short-lived sessions: log in (by a human or a secrets-managed flow) at the start, expire at the end.
  • Separate accounts per client in agency settings, so one client's run can never touch another's data.
  • Keep an inventory: which agent accounts exist, what they can do, who owns them, when they were last reviewed.

Worked example: a creator-management agency in the UK

A London talent agency wanted an agent to download monthly analytics screenshots from creators' platform dashboards. The first idea was to reuse each creator's login in a shared browser. Instead they (1) used official analytics APIs or exports wherever the platforms offered them, (2) for the remaining dashboards, had creators grant delegated, read-only access where the platform supports it, (3) ran each creator's task in a fresh container with an allowlist containing only that platform's domains, and (4) never stored creators' passwords. When the agent misbehaved in testing (it tried to open a messages tab), the role simply did not allow it.

Pitfalls

  • Running agents in your personal browser "just to try it" and forgetting to stop.
  • Allowlisting whole categories ("any .com") instead of specific domains.
  • One shared super-account for all automation.
  • Persistent environments that accumulate cookies and downloads across clients.

How to measure success

Audit monthly: number of agent accounts and their roles, share of runs in fresh environments, blocked egress attempts (a spike is a signal of injection or bugs), and time to revoke an agent's access.

Key takeaways

  • Assume the agent will err or be manipulated; isolation and least privilege make that moment harmless.
  • Isolate compute, browser profile, network, identity and data; destroy or reset environments after each run.
  • Use an egress proxy allowlist plus harness-level domain checks as two independent controls.
  • Give agents dedicated, minimal, preferably read-only accounts, separated per client, with an inventory.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which setup best limits the blast radius of a browser agent?
  2. Why enforce the domain allowlist in both the proxy and the harness?
  3. An agency automates tasks for ten clients. What account design is best?

Put it into practice

Write the isolation spec for one agent workflow: environment type, allowed domains, account role, data inputs, and how the environment is reset after each run.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.