arXiv cs.CR·2d agoAEGIS: Audio Endogenous Guarding via Internal Signals Against Large Audio-Language Model Jailbreaks#aegis#jailbreak#audioAI safety & security
arXiv cs.AI / cs.LG / cs.CL·8d agoWhen Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence#vision-language-models#human-robot#prompt-sensitivityAI research
arXiv cs.CR·18d agoHow Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE#alignment#directional-ablation#glmAI safety & security
Hugging Face Blog·18d agoSafety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic#ai-safety#alignment#model-behaviorAI safety & security
Hugging Face daily papers·24d agoSafety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal#llm#over-refusal#qwen3-8b1
Hugging Face daily papers·29d agoRecognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions#abstention#alignment#interpretability