Google Launches Gemini 4 Argon With Guardrail-Free Access for Vetted Defenders
Google launches frontier model Gemini 4 Argon, initially without cyber guardrails for vetted defenders.
Google announced Gemini 4 Argon, a frontier model initially offered to trusted cyber defenders and internal teams through the Fairwind Program without cyber guardrails. Google says Argon can autonomously find, validate, and patch vulnerabilities and uncovered an unnamed critical flaw exposing personal data in hospital software used worldwide. On Collinear AI's CWE-bench v1, Argon tied for first at 68% with OpenAI's GPT-6 Astra and xAI's Grok 4.7. A broader release is intended to refuse cyber and CBRN misuse, monitor chain-of-thought for misalignment, and resist indirect prompt injection.
- Limited Fairwind access removes cyber guardrails for vetted defenders.
- Argon reportedly found a critical vulnerability in unnamed hospital software.
- CWE-bench v1 score of 68% ties GPT-6 Astra and Grok 4.7.
- Wider release adds CBRN refusals and prompt-injection mitigations.
Full article579 words · extracted from securityweek.com · click to collapse
Google on Wednesday announced Gemini 4 Argon, its new frontier AI model, which it is initially rolling out to a select group of trusted cyber defenders through its Fairwind Program.
Google’s internal teams are also using the model. The company says it will expand access gradually as it gathers feedback from early testers.
The tech giant says Argon is designed for complex workflows in software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. On the security side, the company says it trained the model to be highly capable at cyber defense, and that it can autonomously find, validate, and patch critical software vulnerabilities.
“For trusted defenders and our own internal teams at Google, we’ll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities,” the company said.
Google launched Fairwind in early September as a limited access program for governments, Google Cloud customers, and cybersecurity partners. It initially combined the Gemini 3.8 Flash Cyber model with Google’s CodeMender harness, which finds, verifies, and fixes vulnerabilities. At launch, the program had more than 650 participating partners.
Google says Argon found a flaw in hospital software
Wiz, which Google acquired earlier this year, is using Argon in its Scan for Good initiative, which finds and remediates high-risk exposures in critical public infrastructure for free.
Advertisement. Scroll to continue reading.
“In an early demonstration of its impact, the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, identifying a severe risk that previous frontier models had missed,” Google said.
The announcement does not name the affected software or say whether the issue has been addressed.
On CWE-bench v1, a vulnerability remediation benchmark developed by Collinear AI, Argon tied for first place with a score of 68%, alongside OpenAI’s GPT-6 Astra and xAI’s Grok 4.7.
The company also claims improvements in vulnerability discovery compared to Gemini 3.8 Flash Cyber. On Google’s internal vulnerability benchmark, Argon found a wide range of exposures across complex codebases written in 20 programming languages.
On Wiz’s internal black-box penetration testing benchmark, which targets live web systems without access to source code, Argon outperformed Gemini 3.8 Flash Cyber. Google says it was better at discovering the attack surface, identifying vulnerabilities, and producing proof-of-concept evidence.
Google is strengthening safeguards before a wider release
“Safely releasing frontier capabilities at this level requires a phased approach,” Google said, noting that it is taking part in the US government’s voluntary process for pre-release model access.
For the broader rollout, Google says the model is designed to refuse requests that could enable cyber or chemical, biological, radiological, and nuclear (CBRN) attacks, while still supporting legitimate dual-use scientific research.
These safeguards include improved techniques for monitoring the model’s internal activations to spot misuse. Google also describes Argon as its most resilient model yet against indirect prompt injection.
“In order to prevent Argon from stepping out of bounds to try to accomplish a task in a way that goes beyond the user’s intentions, we are deploying misalignment mitigations that monitor Argon’s chain-of-thought and actions and stop execution when necessary,” the company explained.
The company is also isolating and sealing its sandboxed environments before high-risk training or evaluations begin.
Related: Google: AI Is Changing the Pace and Profile of Vulnerability Discovery
Related: Anthropic Flags AI Agent Liability Risks as OpenAI Faces Hacking Lawsuit
Related: Trump Says Top Tech Firms Have Signed Accord to ‘Self-Police’ AI Development