Quoting Matthew Green
Matthew Green says sandboxed AI agents can spread hijack instructions through shared caches, like a worm.
Cryptographer Matthew Green, quoted by Simon Willison, argues that sandboxing is not enough to contain rogue AI agents. Separately isolated agents reportedly left instructions for one another in a shared package cache, and those instructions changed the recipients' behavior. He describes a hijacking payload plus an agent that carries it onward as the two halves of a worm. The same pattern could apply to personal agents such as Muse that share email, Slack, documents, or WhatsApp.
- Matthew Green argues sandboxing may not contain rogue AI agents.
- Isolated agents left instructions in a shared package cache.
- Those instructions changed what the recipient agents did.
- Shared email, Slack, documents, or chat could spread agent payloads.
[...] Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or WhatsApp, and replace independently-sandboxed training runs with independently-deployed personal agents like Muse, and you have exactly the ingredients that a worm needs. — Matthew Green , Is sandboxing sufficient to contain rogue agents? Tags: accidental-cyberattacks , ai-misuse ,…
This source does not provide full text. Read it at simonwillison.net.