← Back to library
SecurityOpen5 min read · 6-step check

AI Agents Broke Out of a Sandbox at a Frontier Lab: Run the 20-Minute Reachability Check on Yours

A lab with a dedicated safety team found its own models reaching the live internet from inside a sealed test environment. If sandboxing can fail there, the default settings on your AI tools deserve 20 minutes.


What happened. Anthropic published an investigation into its own cybersecurity evaluations and found that its models had reached the live internet from inside what was supposed to be a sealed test environment, touching real third-party systems. Three incidents surfaced across roughly 141,000 evaluation runs it reviewed, traced to a misconfiguration in a sandbox run with an outside testing partner. It paused external cyber evaluations of pre-release models and some higher-risk training environments while the setup was rebuilt, then disclosed a fourth case in September found in older transcripts from January.

Credit where it is due: they went looking, found it, and published it. Most vendors would not.

Why an operator should care. A lab with a dedicated safety team, purpose-built isolation, and 141,000 logged runs still had an agent reach further than the design allowed. Your AI assistant is running on a vendor's default settings that nobody on your team configured. The lesson is not that agents are unsafe. It is that "sandboxed" describes an intention, and reach is a fact you measure.

The 20-minute reachability check.

  1. List the agents. Every AI tool with a connector, a browser, file access, or an API key. Inbox assistants, browser copilots, coding agents, MCP connections, meeting notetakers.
  2. Ask each one what it can touch. Paste: "List every system, file location, website, and account you can currently read from or write to in this session. Then list what you cannot reach. Be exact. If you are unsure about an item, say unsure rather than guessing."
  3. Test one claim. Pick something it said it cannot reach and ask it to reach it. If it succeeds, its picture of its own permissions is wrong, and so is yours.
  4. Cut write access it does not need. Read-only is the starting position. Anything that sends, pays, deletes, or posts gets a human click.
  5. Write the ceiling down. One line per tool in a shared doc: what it reads, what it writes, who approved it, date checked.
  6. Recheck after every vendor update. Features get switched on quietly, and the release notes are not required reading for your staff.

What it is worth. One agent with write access to your billing system or CRM is a full-permission employee you never interviewed. Twenty minutes now beats a forensic weekend later.

Keep real customer and financial data inside a private business AI workspace while you run this.