Hugging Face daily papers·6d ago1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation#on-policy-distillation#gradient-estimation#token-selection
arXiv cs.AI / cs.LG / cs.CL·9d agoScore Centering Stabilizes Off-policy Reinforcement Learning#llm-training#quantization#reinforcement-learningAI research1
Hugging Face daily papers·13d agoMind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States#distillation#human-aware-ai#llm-training1
arXiv cs.AI / cs.LG / cs.CL·15d agoExpert-Space Exploration in MoE Reinforcement Learning#expert-routing#grpo#llm-trainingAI research
Hugging Face daily papers·16d agoLearning to Solve Hard Problems in RL for LLMs by Never Giving Up#adaptive-sampling#coding-benchmark#grpo
Hacker News · security·17d agoTraining a 3.8B LLM to 0.384 CORE for $998 – Hugo Vergnes#core-benchmark#fp8#little-lmAI research 15 min1
arXiv cs.AI / cs.LG / cs.CL·18d agoToolLoop: Closed-Loop Tool-Use Data Synthesis via Decomposed Generation and Dynamic Self-Feedback#acebench#bfcl#function-callingAI research1
Hugging Face daily papers·19d agoDifficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR#llm-training#pass-at-k#reasoning2
Latent Space·Aug 22, 2026[AINews] 10% worse, 100x cheaper, 10000x faster: Why Simulation is taking over#autoresearch#distillation#glm 15 min1