arXiv cs.CR·10d agoSafety Beyond the Interface: Detecting Harm via Latent States in Large Language Models#guardrails#harm-detection#llama-3.1AI safety & security1