TL;DR: two news items about isolating agents:

  • anthropic published documentation on configuring auto mode: the classifier checks the tool calls that the ordinary allow, deny and ask rules did not resolve, and the autoMode block helps configure it — you describe your environment there and add your own rules on top of the built-in ones.
  • docker released sbx — isolated micro-vms where the agent gets its own kernel and its own docker daemon, while outbound traffic goes through a host-side proxy with a list of allowed domains.

The classifier understands intent, the sandbox keeps the agent from reaching. Both are worth running.

The classifier inside the process

Auto mode removes routine confirmations: instead of a question on every command, the classifier itself checks the calls that the ordinary allow, deny and ask rules did not resolve. It blocks irreversible and destructive actions, along with anything reaching beyond the trusted environment. The default trust boundary is the working directory and the configured remote of the current repository. Everything else counts as a possible leak destination until you name it explicitly. Though this doesn’t always fire ¯\_(ツ)_/¯

The configuration is four arrays of rules, written as text rather than regexes and tool masks:

  • environment — a description of the environment: which host the repositories live on, which domains and buckets are yours, which services are internal, what counts as sensitive
  • allow — exceptions to the soft_deny rules
  • soft_deny — destructive actions, overridable by explicit intent
  • hard_deny — an unconditional ban, nothing overrides it

There are many built-in rules in soft_deny, and external destinations are only part of them. They cover deletion along an unverified path, force push, deploying to production, disabling tls verification, writing a secret to a file, merging without review, approving your own pr, deleting servers, editing the agent’s own permissions, actions through chrome-mcp. Crossing the trust boundary is handled by the single hard_deny rule: it forbids data leaving the boundary.

Order of precedence: hard_deny → soft_deny → allow → the user’s direct intent. Intent counts only when the message describes that exact action: “clean up the repository” does not authorise a force push, “force-push this branch” does.

A few more details you don’t notice in the documentation right away:

  • the literal "$defaults". It splices the built-in list in place, and without it your array replaces the built-in rules entirely. Set soft_deny without "$defaults" and you lose the whole built-in list
  • autoMode is read only from user and managed settings. Not from .claude/settings.json or .claude/settings.local.json: both live in the repository directory, so a repository could grant itself permissions
  • narrow rules like Bash(npm test) resolve before the classifier, so the prefix can let through an argument nobody anticipated. Cured by the classifyAllShell: true flag — while auto mode is active it suspends every allow rule for Bash and PowerShell, so the classifier checks each command
  • rules from all scopes add up: a developer can extend any list but cannot remove a managed entry. Meanwhile their own allow neutralises an organisational soft_deny. The firm boundary is only permissions.deny in managed settings, which resolves before the classifier

From 14 August 2026 auto mode becomes the default permission mode for new sessions on the Pro, Max and Team plans.

The micro-vm sandbox

docker solves the same problem, but with a hypervisor. Each sandbox is a separate micro-vm: it boots in seconds and lives until you remove it. docker desktop is not needed, sbx itself is free — only centralised management costs money.

Isolation is assembled from five layers.

  • Hypervisor. Its own kernel per sandbox, no shared memory and no shared processes with the host.
  • Network. All http and https goes through a proxy on the host. Everything is denied by default except a list of domains, and the filter sits at dns too: a name outside the list simply gets no address back. Raw tcp is cut not at the connection but at the data: connect succeeds, then nothing flows. udp and icmp are closed, the host’s localhost is unreachable.
  • Docker. Its own daemon and its own image cache inside, no layer reuse between sandboxes.
  • Working directory. Mounted at the same absolute path as on the host, so paths in build logs look the way they normally do. The --clone flag hands over the repository read-only, and the agent works with a private copy.
  • Credentials. api keys never get inside: the Authorization header is supplied by the proxy on the host. A request from the sandbox with a fake token gets a 200 where the same request from the host gets a 401.

What covers what

The layers are independent, and the difference shows in what each one lets through.

The classifier works with intent. It understands that sending repository contents to a third-party api is a leak risk, even when technically it is an ordinary curl. But it lives in the same process as the agent, and its rules are instructions interpreted by a model. That protection is simpler and more flexible, but non-deterministic and evadable.

The sandbox works with mechanics. It does not care what the agent intended — it simply does not let it reach the host or the domains that aren’t allowed. But it is harder to set up and demands more attention while you do.

The shared weakness of both approaches is the lists of permissions and trusted destinations that you have to maintain by hand. Then again, agents can be put to that job :D