SecurityWeek·12h agoChina and US Agree to Establish AI Safety Channel and Continue Trade and Military Talks#ai-policy#us-china#diplomacy 3 min
The Verge · AI·14h ago highOpenAI pauses training of its ‘most capable models’#openai#ai-safety#agent-security 2 sources in the wild2
The Decoder·1d ago highOpenAI's agents went after government and university sites months before Hugging Face#openai#ai-agents#ai-safety 22 sources in the wild 5 min
The Decoder·1d agoAnother Google Deepmind researcher quits, says building superintelligent AI soon is "inherently irresponsible"#google-deepmind#ai-safety#superintelligence
The Verge · AI·1d ago highOne company is at the center of a wave of rogue AI attacks#irregular#ai-agents#ai-safety in the wild 4 min
The Hacker News·2d agoAnthropic and OpenAI Models Still Attempt Restricted Actions in Safety Tests#ai-models#ai-safety#alignment 12 sources 4 min
TechCrunch · AI·2d agoIs the AI industry really ready to slow down?#ai-policy#ai-safety#anthropic 3 sources 7 min
arXiv cs.AI / cs.LG / cs.CL·2d agoJevOut: Natural Context Can Flip Decision Models#decision-models#jev#robustnessAI safety & security
arXiv cs.CR·2d agoInstrumental Monitor Evasion Emerges Under Ordinary Task Pressure#ai-safety#evasionbench#monitor-evasionAI safety & security
arXiv cs.CR·2d agoENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation#prompt-injection#endoprompt#llm-securityAI safety & security
TechCrunch · AI·2d agoFive AI safety sessions every founder should have on their TechCrunch Disrupt 2026 agenda#techcrunch#disrupt#conference 2 sources 5 min
The Verge · AI·2d agoWhy can’t we just keep rogue AIs off the internet?#ai-safety#air-gap#ai-agents 5 min
arXiv cs.CR·2d ago highHard Stop: Kernel-Level Preemption and Containment for Rogue Agentic Execution#agentic-ai#sandbox-escape#hugging-faceAI safety & security
Hugging Face daily papers·2d agoJust Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures#ai-safety#alignment#jailbreak 2 sources
arXiv cs.CR·2d agoAEGIS: Audio Endogenous Guarding via Internal Signals Against Large Audio-Language Model Jailbreaks#aegis#jailbreak#audioAI safety & security
Ars Technica · AI·3d agoTrump’s China rivalry and “AI race” delusion may endanger US, experts say#ai-policy#us-china#ai-safety 2 sources 9 min
MarkTechPost·3d agoAnthropic Releases Claude Opus 5.5: Fable 5.1-Level Performance at 40% Lower Running Cost Than Opus 5#anthropic#claude#opus-5.5 9 sources 5 min
SecurityWeek·3d agoWorries About an AI Internet Takeover Gain New Urgency Among Doomsday Scenarios#ai-safety#ai-agents#openai in the wild 5 min
arXiv cs.AI / cs.LG / cs.CL·3d agoAn Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice#eu-ai-act#systemic-risk#ai-evalsAI safety & security
SecurityWeek·3d agoOuterlimit Raises $16 Million to Stop Rogue AI Agents From Causing Harm#agentic-ai#ai-safety#alignment 2 sources 4 min
arXiv cs.CR·3d ago highYour Model Is Leaking: Covert Information Transfer through LLM Residual Streams#ai-safety#covert-channel#llmAI safety & security in the wild
CyberScoop·3d agoThe president has called for AI leadership. Here’s the mission.#ai-policy#ai-safety#critical-infrastructure 8 min