Hugging Face daily papers·3d agoYour Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs#transformer#superposition#llm
Hugging Face daily papers·3d agoParts-of-Speech as Emergent Categories in SAE Latent Space#sparse-autoencoders#interpretability#part-of-speech
arXiv cs.AI / cs.LG / cs.CL·3d agoOrder-Invariant Answers, Order-Sensitive Representations in Mathematical Reasoning#language-models#mathematical-reasoning#representationsAI research
Hugging Face daily papers·4d agoCapable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models#chain-of-thought#reasoning#frontier-models 2 sources
Hugging Face daily papers·5d agoThe Linear Representation Hypothesis Needs a Group Action#interpretability#linear-representation-hypothesis#group-actions
arXiv cs.AI / cs.LG / cs.CL·5d agoLinguistic Features for Interpretable Textual Entailment#nlp#textual-entailment#interpretabilityAI research
Hugging Face daily papers·7d agoWhy Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms#video-diffusion#interpretability#rope
arXiv cs.AI / cs.LG / cs.CL·8d agoDiaVLo: Diagnosing Behaviours of Vision-Language Models#vision-language-models#diagnostics#alignmentAI research
arXiv cs.AI / cs.LG / cs.CL·8d agoA Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal#sandbagging#unlearning#deceptionAI safety & security
The Decoder·8d agoVisible chains of thought are a safety advantage for AI, but that transparency is slipping away#ai-safety#chain-of-thought#gemini-3-pro1
Hugging Face daily papers·9d agoA Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal#alignment#deception#evals
arXiv cs.AI / cs.LG / cs.CL·9d agoSemantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights#agents#human-ai-interaction#interpretabilityAI research
TechCrunch · AI·9d agoBase Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire#abliteration#ai-safety#base-labs 2 min
arXiv cs.AI / cs.LG / cs.CL·9d agoNeuSOGA3D: A Neuro-Symbolic Framework for Explainable 3D Geometric Reconstruction#3d-reconstruction#cad#interpretabilityAI research
arXiv cs.AI / cs.LG / cs.CL·10d agoFlag Game: A Toy Model for Mechanistic Swarm Interpretability#alignment#belief-formation#emergent-behaviorAI safety & security
arXiv cs.AI / cs.LG / cs.CL·10d agoMonitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations#alignment#frontier-models#interpretabilityAI safety & security1
arXiv cs.AI / cs.LG / cs.CL·10d agoProbabilistic Linear Explanations#explainability#formal-methods#interpretabilityAI research
arXiv cs.AI / cs.LG / cs.CL·11d agoLarge Language Models Develop Belief State Geometry In-Context#belief-state#hmm#in-context-learningAI research1
arXiv cs.AI / cs.LG / cs.CL·12d agoA Chosen Future Can Still Be Rewritten: Causal Writability in Video Models#interpretability#model-editing#physics-simulationAI research
arXiv cs.AI / cs.LG / cs.CL·12d agoDisentangling Representation Evolution in Transformers through Directional Decomposition#interpretability#model-editing#pretrainingAI research1
Hugging Face daily papers·13d agoMoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup#efficient-inference#interpretability#memory-embeddings2
Hugging Face daily papers·13d agoDisentangling Representation Evolution in Transformers through Directional Decomposition#compression#interpretability#model-editing
The Decoder·14d agoAI models' written reasoning steps correspond to distinct internal patterns, a new study finds#activation-analysis#chain-of-thought#interpretability 4 min2
arXiv cs.AI / cs.LG / cs.CL·15d agoMAxBench: A Multinomial Concept Recovery Benchmark#benchmark#concept-recovery#evaluationAI research