arXiv cs.AI / cs.LG / cs.CL·2d agoJevOut: Natural Context Can Flip Decision Models#decision-models#jev#robustnessAI safety & security
arXiv cs.CR·2d agoJust Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures#alignment#evals#benchmarkAI safety & security 2 sources
Security Affairs·7d agoGoogle Gemini also Broke Out of Its Test Environment#agentic-ai#ai-safety#containment 4 min
The Decoder·7d agoGPT-6 Astra and Claude Fable turn robot arms into slapstick killer robots in new safety benchmark#benchmark#claude-fable-5.1#evals 3 min
The Decoder·7d agoGoogle's Gemini also accidentally hacked three real companies during security testing#ai-agents#ai-safety#evals 2 min
arXiv cs.CR·8d agoAPort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport#ai-agents#agent-security#payment-authorizationAI safety & security
arXiv cs.AI / cs.LG / cs.CL·8d agoA Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal#sandbagging#unlearning#deceptionAI safety & security
arXiv cs.AI / cs.LG / cs.CL·8d agoWhen Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence#vision-language-models#human-robot#prompt-sensitivityAI research
Hugging Face daily papers·9d agoA Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal#alignment#deception#evals
Hacker News · AI·10d agoDeepSeek v4.1 Flash Is Now Our Best Hacking Model#ai-agents#autonomous-exploitation#deepseek 5 min
Hugging Face daily papers·11d agoPACT: Can Enterprise AI Assistants Be Trusted Under Pressure?#alignment#benchmark#compliance
TechCrunch · AI·11d agoEarly Anthropic hire, former METR COO have found a way to rein in rogue AI agents#agent-security#ai-safety#aiuc 3 min
Latent Space·12d ago[AINews] AEF-1 standard emerges for Third Party Evaluators, as Xai, OpenAI, and Anthropic all cosign#ai-safety#anthropic#evals 10 min1
arXiv cs.CR·12d agoApproval Integrity and Recovery in LLM Answer Publication#authorization#evals#llm-securityAI safety & security
Hacker News · AI·13d agoClaude Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher#ai-capability#claude#cryptanalysis 7 min1
MarkTechPost·15d agoAnthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills#anthropic#ci#claude-code 3 min2
TechCrunch · AI·16d agoAnthropic reveals rogue AI agents hate CAPTCHAs, just like you#agentic-misbehavior#agent-security#anthropic 5 min1
Cyber Security News·17d agoClaude AI Models Gained Unauthorized Access to Real Systems During Cybersecurity Tests#agentic-ai#ai-safety#alignment 4 min1
arXiv cs.AI / cs.LG / cs.CL·17d agoCross-Model Agreement as a Deployment-Time Reliability Signal for Automatic Polyp Segmentation#cross-model-agreement#evals#medical-imagingAI research
arXiv cs.AI / cs.LG / cs.CL·18d agoSAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?#ai-agents#benchmark#evalsAI research2
Hacker News · security·19d agoHow well do agents use test/verification techniques?#ai-agents#codex#evalsAI research 15 min
Hugging Face daily papers·19d agoSAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?#ai-agents#benchmark#evals1