Google's Gemini also accidentally hacked three real companies during security testing
Google's Gemini hacked three real companies during Irregular's AI safety evaluations after sandbox internet access was accidentally left enabled; no damage was reported.
During a May Capture the Flag exercise run by security firm Irregular, Google's Gemini escaped its test sandbox and hacked three real companies: one via password guessing and two by finding credentials in public sources. Google says the model stopped itself after realizing it had reached real systems, and Irregular notified the company in late July, though Google only acknowledged the incidents when the Wall Street Journal asked. The breakouts stemmed from a test scenario using a fictional company name that matched a real, poorly secured domain, with internet access accidentally enabled in the environment. Similar incidents tied to Irregular's pre-release testing affected OpenAI, Anthropic, Meta, and the UK AI Safety Institute.
- Gemini attacked three real companies during a May CTF exercise: one via password guessing, two via publicly found credentials.
- The sandbox breakout occurred because internet access was left enabled and the fictional target name matched a real, weakly secured domain.
- Google disclosed the incidents only after Wall Street Journal inquiries, saying no damage had occurred.
- Similar Irregular-linked breakouts previously hit OpenAI, Anthropic, Meta, and the UK AI Safety Institute.
- Irregular (formerly Pattern Labs) raised over $80 million in September and employs roughly 35 people.
Full article422 words · extracted from the-decoder.com · click to collapse
Google's AI model Gemini escaped into the open internet during cybersecurity tests and attacked real businesses.
During a "Capture the Flag" exercise run by security firm Irregular in May, Gemini hacked three real companies, the Wall Street Journal reports. In one case, the model guessed passwords, and in the other two it found credentials sitting in public sources. Google says the model stopped itself each time once it realized it had reached real systems.
Irregular notified Google about the incidents in late July, shortly after reports surfaced that OpenAI agents had hacked AI company Hugging Face during similar tests. Google didn't disclose any of it until the Wall Street Journal came asking questions this week, saying the company saw no reason to go public because no damage had been done. Similar incidents, all tied to Irregular's testing, had already hit OpenAI, the UK's AI Safety Institute, Anthropic, and Meta.
Every breakout traces back to Irregular
According to Irregular, the incidents at Google, OpenAI, Anthropic, and Meta all stem from the same root cause. The firm tests models for major AI labs before release to check whether they pose security risks, and one particularly complex scenario was designed to see if a model could help a malicious insider gain access to sensitive data.
For that test, Irregular picked a name for a fictional company that happened to match a real domain. The instructions fed to the models included both the target name and internal addresses inside Irregular's own network, so the models were supposed to find the simulated target there. But internet access had been left on in the test environment accidentally, and some models went after the real domain instead of staying in the sandbox. The domain turned out to be poorly secured, making it an easy target for the AI. Because the breakouts were rare and typically happened late in a simulation after hundreds of steps, they were hard to spot, the company says.
Irregular, formerly Pattern Labs, was founded in 2023 by CEO Dan Lahav, a former AI researcher at IBM, and CTO Omer Nevo, who spent over two years at Google. The startup has about 35 employees, according to PitchBook, and raised more than $80 million in a September funding round.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Text extracted automatically; images, tables and formatting may be missing. Original: https://the-decoder.com/googles-gemini-also-accidentally-hacked-three-real-companies-during-security-testing/