arXiv cs.CR·2d agoPrefilling the Reasoning Channel: Output-Prefix Attacks on Reasoning LLMs#prompt-injection#jailbreak#reasoning-modelsAI safety & security
Hugging Face daily papers·2d agoJust Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures#ai-safety#alignment#jailbreak 2 sources
arXiv cs.CR·2d agoAEGIS: Audio Endogenous Guarding via Internal Signals Against Large Audio-Language Model Jailbreaks#aegis#jailbreak#audioAI safety & security
Schneier on Security·3d agoResearch on Models Engaging in Genie-Like Behavior#ai-safety#self-jailbreaking#jailbreakAI safety & security
arXiv cs.CR·5d ago highDecoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection#adversarial#guardrails#jailbreakAI safety & security
Security Affairs·5d agoThe Target Is No Longer the Model. It’s the Agent.#agent-supply-chain#ai-agents#echoleak 9 min
arXiv cs.CR·8d agoCASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation#jailbreak#llm-security#defense-in-depthAI safety & security
arXiv cs.CR·8d agoHE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference#homomorphic-encryption#jailbreak#llmAI safety & security
Help Net Security·10d agoOne runaway AI agent racked up a $50,000 cloud bill#ai-agents#ai-supply-chain#gtig in the wild 4 min3
Help Net Security·11d agoUncensored AI sold on hacking forum as alternative to ChatGPT and Claude jailbreaks#cybercrime#jailbreak#malicious-llm 2 min
arXiv cs.CR·12d agoDivide, Consult, Conquer: Capability Laundering Through Aligned LLMs#capability-laundering#cbrn#frontier-modelsAI safety & security
GBHackers·16d agoNew AI Attack Hides Malicious Instructions in Normal-Looking Text to Evade Safety Filters#agents#ai-safety#checkpoint 5 min
Check Point Research·16d agoPuzzleMask: Abusing Plain Prose as a Covert AI Attack Vector#ai-safety#check-point#gpt-5 15 min
Ars Technica · AI·17d agoSix Chinese AI firms accused of aggressively copying US frontier models#api-abuse#cisa#deepseek in the wild 8 min1
arXiv cs.CR·18d agoCS-Guard: Benchmarking LLM Guardrails for Code Generation Security#benchmark#code-generation#cs-guardAI safety & security1
Schneier on Security·18d agoStealing AI Reasoning Traces#chain-of-thought#distillation#encryptionAI safety & security 2 min
arXiv cs.CR·18d agoStructural Jailbreaks Generalize but Do Not Compound: A cross-provider and multilingual study of Involuntary In-Context Learning#benchmark#gemini#iiclAI safety & security1
The Register · Security·19d agoOpenAI's rebel agent swarm died young, but its chilling logs live on#agent-safety#ai-agents#alignment 5 min
Security Affairs·Aug 23, 2026Zero-Click Grok Chat History Theft: Adversa AI Demonstrates Cryptographic Context Injection#adversa-ai#agent-security#cryptography 5 min
Palo Alto Unit 42·Aug 17, 2026Bad Likert Judge: A Novel Multi-Turn Technique to Jailbreak LLMs by Misusing Their Evaluation Capability#ai-safety#jailbreak#llm 15 min