GBHackers·9h ago highOpenAI Says Misaligned AI Agents Hacked Hugging Face and Bypassed Security Controls#openai#hugging-face#ai-agents 2 sources in the wild 3 min
Ars Technica · AI·1d agoOpenAI agent “didn’t accept no for an answer” in Australian government breach#openai#ai-agents#misalignment 22 sources 5 min
Latent Space·1d ago[AINews] Meta Connect 2026: Muse glasses, voice, video, and Charm#meta#muse#meta-connect 11 sources 14 min
Hugging Face daily papers·5d agoRecursive self-improvement of AI research agents#recursive-self-improvement#ai-agents#aide
Hugging Face daily papers·6d agoRULER: Instance-aware Rubric Rewards for SVG Generation#svg-generation#reinforcement-learning#rubric-rewards
Cyber Security News·9d agoOpenAI Models Searched for Leaked API Keys and Uploaded Files Without Permission#agentic-ai#ai-safety#data-exfiltration 4 min1
Ars Technica · AI·9d agoCovert uploads and megalomania: OpenAI details new "misaligned" agent incidents#agentic-ai#ai-alignment#ai-safety 5 min1
arXiv cs.AI / cs.LG / cs.CL·10d agoMonitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations#alignment#frontier-models#interpretabilityAI safety & security1
Hacker News · AI·11d agoWhy I'm still bearish on LLMs after Navier-Stokes#agents#autonomy#benchmarksAI industry 5 min1
Hugging Face daily papers·12d agoImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals#alignment#benchmark#evaluation
MarkTechPost·13d agoAnthropic’s 3-Step ‘Pace the Frontier’ Plan Wins OpenAI, xAI and Microsoft Support: Is It Too Late to Slow AI Down?#agent-misalignment#ai-safety#anthropic 9 min1
The Verge · AI·15d agoAnthropic spent this week in hot water over cybersecurity#agentic-ai#alignment#anthropic 4 min1
Hugging Face daily papers·16d agoSteerDuplex: Steerable Duplex Speech Dialogue Models#benchmark#full-duplex#moshi
arXiv cs.CR·17d agoBenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure#benchmark-integrity#evaluation-infrastructure#llm-agentsAI safety & security1
Hacker News · security·17d agoA Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming#ai-alignment#deepmind#opinion 8 min
arXiv cs.AI / cs.LG / cs.CL·18d agoEntropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation#arxiv#code-generation#entropy-regularizationAI research1
Hugging Face daily papers·19d agoSWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents#benchmarks#coding-agents#evaluation1
Security Affairs·19d agoWhy AI Agent Sandboxes Are Failing Security Tests#agentic-ai#agent-security#ai-safety in the wild 6 min
MarkTechPost·20d agoIFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B#apache-2.0#ifm#k2-horizon 4 min1
The Hacker News·Aug 28, 2026 highOpenAI Says Reward Hacking Drove AI Agents to Exploit Zero#agentic-ai#artifactory#hugging-face in the wild 7 min