Hugging Face daily papers·3d agoParts-of-Speech as Emergent Categories in SAE Latent Space#sparse-autoencoders#interpretability#part-of-speech
Hugging Face daily papers·5d agoThe Linear Representation Hypothesis Needs a Group Action#interpretability#linear-representation-hypothesis#group-actions
arXiv cs.AI / cs.LG / cs.CL·9d agoDeep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models#activation-steering#agent-security#llmAI research
arXiv cs.AI / cs.LG / cs.CL·11d agoLarge Language Models Develop Belief State Geometry In-Context#belief-state#hmm#in-context-learningAI research1
arXiv cs.AI / cs.LG / cs.CL·12d agoThe Router Within: Eliciting Native Skill Routing from a Frozen LLM#linear-probe#llm-agents#mechanistic-interpretabilityAI research1
Hugging Face daily papers·13d agoDecoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration#jailbreak-defense#llama-3#llm-safety
arXiv cs.AI / cs.LG / cs.CL·16d agoFrom Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge#gemma#hidden-states#interpretabilityAI research1
arXiv cs.AI / cs.LG / cs.CL·18d agoSAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?#ai-agents#benchmark#evalsAI research2
arXiv cs.AI / cs.LG / cs.CL·19d agoLLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders#ai-security#language-models#llm-backdoorsAI safety & security1
Hugging Face daily papers·23d agoBeneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs#chain-of-thought#interpretability#llm1