Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity
Anthropic launched Claude Opus 5.5 with cybersecurity and biology safeguards that reroute risky requests to weaker models.
Anthropic released Claude Opus 5.5, its first model since CEO Dario Amodei said the company would slow frontier development. The company says it is cheaper and more efficient than Opus 5, matches Claude Fable 5.1 on most work, and scores best on Anthropic's broadest alignment test. Cybersecurity requests are rerouted to Opus 4.8 and flagged biology requests to Opus 5, using safeguards similar to Fable 5.1. Frontier Design and METR tested it before release, and Sonnet 5.5 and Haiku 5.5 are planned in the coming weeks.
- First Anthropic release after Dario Amodei pledged to pace frontier development.
- Cybersecurity requests route to Opus 4.8; flagged biology requests route to Opus 5.
- Anthropic calls it the strongest model on its broadest alignment test.
- Frontier Design and METR tested the model before release.
- Sonnet 5.5 and Haiku 5.5 are planned in the coming weeks.
Full article260 words · extracted from theverge.com · click to collapse
Emma Roth
is a news writer who covers the streaming wars, consumer tech, crypto, social media, and much more. Previously, she was a writer and editor at MUO.
Anthropic says its new Claude Opus 5.5 model comes with stronger safeguards in the wake of recent rogue AI hacking incidents. In an announcement on Tuesday, Anthropic says Opus 5.5 comes with improvements to certain risky behaviors, including attempts to escape the company’s testing sandbox.
It’s the first model released by Anthropic after CEO Dario Amodei announced plans to “pace the frontier,” or slow down AI development. In recent weeks, several AI companies, including Anthropic, Google, and OpenAI, have reported that their AI models escaped containment and hacked third-party companies during testing.
Anthropic says Opus 5.5 is the “strongest performing” model on the company’s most comprehensive alignment test. The model, which is cheaper and more efficient to run than Opus 5, will come with safeguards similar to the ones offered by Anthropic’s more advanced Fable 5.1 model. That means Opus 5.5 will re-route certain cybersecurity-related requests to the less powerful Opus 4.8, while biology-related requests flagged by its safeguards will go to Opus 5.
Opus 5.5 also matches the performance of Fable 5.1 “on most work,” and was tested by outside partners, including Frontier Design and METR, before release. The company also plans to launch Claude Sonnet 5.5 and Haiku 5.5 in the coming weeks.
Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.
- Emma Roth
Text extracted automatically; images, tables and formatting may be missing. Original: https://www.theverge.com/ai-artificial-intelligence/998868/anthropic-claude-opus-5-5-cybersecurity