The Decoder·14h agoAI access makes people almost entirely unwilling to say "I don't know," study finds#ai-overreliance#human-judgment#epistemia 4 min
arXiv cs.AI / cs.LG / cs.CL·2d agoMinimally Invasive Steering of Language Models#language-models#steering#misvoAI research
Hacker News · AI·3d agoContrastive Language Models#language-models#contrastive-learning#hacker-newsAI research
Hugging Face daily papers·3d agoParts-of-Speech as Emergent Categories in SAE Latent Space#sparse-autoencoders#interpretability#part-of-speech
arXiv cs.AI / cs.LG / cs.CL·3d agoOrder-Invariant Answers, Order-Sensitive Representations in Mathematical Reasoning#language-models#mathematical-reasoning#representationsAI research
Hugging Face daily papers·4d agoCapable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models#chain-of-thought#reasoning#frontier-models 2 sources
Hugging Face daily papers·5d agoMemoryAthena: Adaptive Routing over Latent and Generated Memories#memoryathena#memory#routing
arXiv cs.AI / cs.LG / cs.CL·8d agoAbstention and Noise Filtering: Two Missing Primitives of Softmax Attention#attention#softmax#gatingAI research
arXiv cs.AI / cs.LG / cs.CL·8d agoA Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal#sandbagging#unlearning#deceptionAI safety & security
arXiv cs.AI / cs.LG / cs.CL·10d agoObjective vs. Search: Decomposing What Makes a Good Tokeniser#benchmark#bits-per-byte#bpeAI research
arXiv cs.AI / cs.LG / cs.CL·15d agoMAxBench: A Multinomial Concept Recovery Benchmark#benchmark#concept-recovery#evaluationAI research
Hugging Face daily papers·17d agoCompetence-Gated Pooling of Language Models and Priors for Event Forecasting#abstention#brier-score#calibration1
arXiv cs.AI / cs.LG / cs.CL·19d agoDecomposition-Guided Diffusion Language Models for Inertial Confinement Fusion Prediction#diffusion-models#icfbench#inertial-confinement-fusionAI research
arXiv cs.AI / cs.LG / cs.CL·19d agoLLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders#ai-security#language-models#llm-backdoorsAI safety & security1
Hugging Face daily papers·20d agoKalman Delta Networks: Uncertainty-aware Associative Memory#associative-memory#kalman-filter#language-models
Hugging Face daily papers·Aug 25, 2026One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation#collapse#distillation#language-models