The Hacker News·21h agoZero Trust for AI Agents Starts With Fixing Zero Visibility#ai-agents#zero-trust#shadow-ai 8 min
The Verge · AI·9d ago highInside the suddenly explosive world of AI safety#agentic-ai#ai-safety#alignment 15 min1
TechCrunch · AI·10d agoAnthropic and OpenAI want to embed safety evaluators. Will they really be independent?#alignment#anthropic#apollo-research 7 min
TechCrunch · AI·11d agoAI agents now have a place to snitch#agent-security#ai-agents#ai-safety 3 min2
TechCrunch · AI·11d agoEarly Anthropic hire, former METR COO have found a way to rein in rogue AI agents#agent-security#ai-safety#aiuc 3 min
MarkTechPost·13d agoAnthropic’s 3-Step ‘Pace the Frontier’ Plan Wins OpenAI, xAI and Microsoft Support: Is It Too Late to Slow AI Down?#agent-misalignment#ai-safety#anthropic 9 min1
Hacker News · security·14d ago highAnthropic CEO says AI swarm could 'take over the Internet' in 6-12 months#ai-agents#ai-safety#anthropic 9 min3
The Verge · AI·14d agoAnthropic CEO says it’s time to pump the brakes on AI#ai-governance#ai-safety#anthropic 2 min
TechCrunch · AI·14d agoAnthropic CEO outlines plan to ‘pace the frontier’#ai-governance#ai-safety#anthropic 4 min3
The Verge · AI·15d agoAnthropic spent this week in hot water over cybersecurity#agentic-ai#alignment#anthropic 4 min1
CSO Online·15d ago highAnthropic finds evidence of a fourth AI escaping from containment#ai-safety#anthropic#claude1
SecurityWeek·16d agoWidened Scan Turns Up Fourth Rogue Claude Cyber Incident#agentic-ai#ai-safety#anthropic 3 min1
Infosecurity Magazine·16d agoAnthropic Reveals Yet Another Cybersecurity Incident#agentic-ai#ai-safety#anthropic 3 min
The Hacker News·17d agoAnthropic Discloses Fourth AI Hacking Incident Involving Claude Opus 4.6#agentic-ai#ai-safety#anthropic 5 min1
Cyber Security News·17d agoClaude AI Models Gained Unauthorized Access to Real Systems During Cybersecurity Tests#agentic-ai#ai-safety#alignment 4 min1
Lobsters · security·17d agoAn alignment assessment of recent cybersecurity incidents#ai-safety#alignment#anthropic 15 min4
Security Affairs·19d agoWhy AI Agent Sandboxes Are Failing Security Tests#agentic-ai#agent-security#ai-safety in the wild 6 min
TechCrunch · AI·22d agoOpenAI’s rogue agents keep escaping, with no formal process to investigate them#ai-agents#ai-safety#incident-investigation 4 min
Ars Technica · Security·22d agoOpenAI agents discussed ways to escape their sandbox on public wiki#agent-collusion#ai-agents#ai-safety 2 min1
The Register · Security·25d agoAttacker stole a METR API key, used $600K worth of credits, and no one noticed for weeks#agent-security#ai-infrastructure#api-key-theft in the wild 4 min1
Dark Reading·25d agoAI Model Evaluator METR Hit by Credential Theft, Probing#ai-infrastructure#api-key#credential-theft in the wild
Infosecurity Magazine·25d agoAttackers Steal METR API Key and Burn $600,000 in AI Credits#ai-misuse#api-key-theft#costAI safety & security in the wild