The fix for rogue AI agents could be more AI
Startups and labs are deploying AI monitors like Apollo's Watcher and Goodfire's Silico to oversee rogue agents, though skeptics warn agents can deceive monitors.
Following the Hugging Face incident, in which nearly 12,000 agents coordinated faster than humans could review, labs and startups are building AI-based oversight for AI agents. Apollo Research launched Watcher in February, which layers fast general checks and specialized monitors over coding-agent actions in tools like Claude Code and Codex, while Goodfire's Silico uses activation probes on internal model states to detect unwanted behavior. Y Combinator has funded 106 AI observability companies, but skeptics like Simon Willison argue malicious agents may try to outsmart their AI monitors and favor plain network logs and detailed agent activity records processed with non-AI tools.