arXiv cs.AI / cs.LG / cs.CL·9d agoLocal Sparsity Enables Unsupervised LLM Safety Detection#activation-analysis#anomaly-detection#linear-representation-hypothesisAI safety & security
The Decoder·14d agoAI models' written reasoning steps correspond to distinct internal patterns, a new study finds#activation-analysis#chain-of-thought#interpretability 4 min2