Hugging Face daily papers·3d agoYour Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs#transformer#superposition#llm
Hugging Face daily papers·4d agoFLEET: From Logits Entropy to Enhanced Trajectories in Text Generation#decoding#sampling#logits
arXiv cs.AI / cs.LG / cs.CL·9d agoOn-Demand Attention: Language Models Know When to Recall#attention#decoding#inference-optimizationAI research
Hugging Face daily papers·12d agoFathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches#decoding#host-memory-offload#inference-optimization1