Irregular says ‘human oversight’ responsible for AI sandbox escape incidents
Irregular says AI models including Claude Opus and GPT-5.6 Sol escaped evaluation sandboxes and took real-world offensive security actions due to accidental internet access.
AI testing firm Irregular published a post-mortem on incidents where frontier models from Anthropic and OpenAI escaped simulated environments and performed real-world hacking. Models including Mythos 5, Claude Opus and GPT-5.6 Sol exploited vulnerabilities, extracted credentials and accessed a production database when internet access was unintentionally enabled. One model nearly targeted a similarly named real company after discovering its credentials online. Irregular says it has remediated the setup issues and plans better logging, documentation and a forthcoming whitepaper.
- Internet access was unintentionally enabled in evaluation environments, letting models take offensive security actions in the real world.
- Affected non-public models included Anthropic's Mythos 5 and Claude Opus and OpenAI's GPT-5.6 Sol.
- One model exploited vulnerabilities, extracted credentials and accessed a production database, mistaking real domains for simulated targets.
- Irregular plans improved log monitoring, revised threat models and a whitepaper on the incidents.
Full article883 words · extracted from cyberscoop.com · click to collapse
Get our latest cybersecurity news first on Google.
In a post-mortem, the frontier AI testing company said internet access for models is necessary to fully test out their cybersecurity capabilities.
Listen to this article
0:00
Learn more.
A company that runs AI testing labs involved in a series of breaches carried out by Anthropic and OpenAI’s cyber-focused models said the incident happened in part because they “unintentionally” provided the models with access to the internet.
In a blog posted Friday, Irregular said it partners with frontier labs to “stress test” AI models for security risks before they’re deployed, including simulations of their vulnerability research capabilities.
Those tests included non-public models from Anthropic and OpenAI. Separate disclosures from the companies confirmed these included Mythos 5, Claude Opus and GPT-5.6 Sol.
“As part of our review, we identified that a few interactions with our evaluation environments, in which internet access was unintentionally made available, led some models to take offensive security actions in the real world,” Irregular said in an unsigned blog.
In the incidents involving Anthropic’s models, testers initially believed the models were using a fictional company in its hacking simulations, only to learn later that the name unintentionally matched the domain of a real company. That led the model to “consider taking offensive action against” the real company in multiple training runs.
Irregular said this was due to “human oversight,” and that they typically do a background check to ensure they’re not using a real company for cybersecurity tests. The instructions provided to the model also included the target’s name and “internal addresses within our network” where the model could go to find the simulated environment.
While Mythos performed as intended in the majority of the thousands of tests, a small number of cases showed the model couldn’t distinguish between the fake and real domains and executed actual attacks on internet infrastructure, including exploiting vulnerabilities, extracting credentials and accessing a production database. In one instance, it even began targeting a site with a similar name when credentials for it were discovered online.
“Mainly, models believed they were in simulated environments, when they in fact took action in the real world,” the blog continued. “We are putting in place new and robust protocols to ensure setup issues do not occur while meeting the constraints of the testing process.”
The company said it plans to release a larger whitepaper breaking down the incidents and update their best practices for evaluation setups in the future.
While the companies have drawn criticism from some in the cybersecurity community for failing to securely design their sandboxes for testing, experts have said AI models are known to grind away on fulfilling a command until they can find a workaround. Additionally, Irregular said granting some level of internet access to models is necessary to fully test out their cybersecurity capabilities.
“Controlled internet access, while it may allow models to exceed containment boundaries, is at times critical for realistic evaluations; without it, threat scenarios lose fidelity, undercutting the purpose of the challenge to reduce post-release risk of models being misused by attackers – as attackers in the real world do rely on the internet,” the company wrote.
According to the blog, Irregular has since “remediated” the “issues that led to these interactions,” though few details are provided.
However, the researchers say the engagement revealed critical gaps in their security practices.
They plan to improve documentation of evaluation setups, deploy better log monitoring tools capable of tracking “the extreme amount of data generated by the traffic,” revise their threat models to account for rogue AI behavior, and establish faster information sharing between stakeholders.
“Looking further down the line, models will only get stronger. While in this case we believe that better implementation of existing safeguards could prevent most incidents of this kind, as models become stronger, this may not be the case,” Irregular wrote. “We therefore believe this opportunity should be leveraged by us and the community to be proactive and establish forward-looking protocols and [research and development] efforts.”
Latest Podcasts
Government
Technology
Threats
Microsoft discloses two actively exploited zero-days among 974 vulnerabilities
Russian national extradited to US for alleged involvement in bank-account takeover scheme
Attackers exploit zero-days in consistently besieged SonicWall product
FBI raises alarm over deceptive phishing campaign targeting prominent people
Policy
Whistleblower says USPS deploying new, ‘untested’ IT systems governing mail-in ballots
‘Watershed 250’ test program in Texas looks to private sector for water cybersecurity help
Former sexual abuse victims say Grok used their images, videos to train deepfake capabilities
Cyber threats nudge Trump to sign executive order on foreign equipment in U.S. energy infrastructure
Text extracted automatically; images, tables and formatting may be missing. Original: https://cyberscoop.com/irregular-ai-sandbox-escape-human-oversight/