Hugging Face daily papers·3d agoParts-of-Speech as Emergent Categories in SAE Latent Space#sparse-autoencoders#interpretability#part-of-speech
arXiv cs.AI / cs.LG / cs.CL·9d agoLocal Sparsity Enables Unsupervised LLM Safety Detection#activation-analysis#anomaly-detection#linear-representation-hypothesisAI safety & security
arXiv cs.AI / cs.LG / cs.CL·18d agoSAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?#ai-agents#benchmark#evalsAI research2
Hugging Face daily papers·19d agoSAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?#ai-agents#benchmark#evals1
arXiv cs.AI / cs.LG / cs.CL·19d agoLLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders#ai-security#language-models#llm-backdoorsAI safety & security1