The Hacker News·8d agoGoogle Gemini Broke Into Real Company Systems After Security Test Domain Mix-Up#agent-safety#ai-agents#capture-the-flag 2 min1
arXiv cs.CR·9d agoClashBench: Conflicts Leading Agents to Seize and Harm#agent-safety#ai-agents#benchmarkAI safety & security
Latent Space·10d agoUnderwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC#agent-safety#aiuc#aiuc-1 15 min
Hugging Face daily papers·12d agoEmergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems#agent-safety#alignment#evaluation
MIT Technology Review · AI·12d agoAI agents blew the whistle on their cheating colleagues#agent-safety#alignment#deepmind 6 min1
arXiv cs.CR·17d agoThe Missing Boundary: How Autonomous Agents Lose Control#agent-safety#autonomous-agents#context-managementAI safety & security2
The Register · Security·19d agoOpenAI's rebel agent swarm died young, but its chilling logs live on#agent-safety#ai-agents#alignment 5 min
The Register · Security·22d agoRogue OpenAI agents used dead German web site to communicate in May, months before Hugging Face incident#agentic-ai#agent-safety#emergent-behavior in the wild 4 min
Import AI·26d agoImport AI 471: Why Hugging Face worries me; space mining; FIve Eyes on AI#agent-safety#ai-agents#alignment 12 min