MarkTechPost·1d agoPerplexity Trains Its Computer Agent on Real Mistakes With Hint-Guided Self-Distillation#perplexity#glm-5.2#self-distillation 4 min
arXiv cs.AI / cs.LG / cs.CL·2d agoPoEM: Predicting RL Outcomes from Existing Policies#poem#reinforcement-learning#reward-modelsAI research
Hugging Face daily papers·3d agoRufus-Air: An Open LLM Post-Training Recipe#post-training#sft#reinforcement-learning
Interconnects·4d agoDebating RSI, the US-China Gap, and Jaggedness with JS Denain of Epoch AI#epoch-ai#rsi#us-chinaAI research 15 min
Hugging Face daily papers·6d agoACLArena: Agent Continue Learning in Multi-stage Post-training#continual-learning#agents#lora
arXiv cs.AI / cs.LG / cs.CL·9d agoOPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher#3dgs#autonomous-driving#end-to-end-drivingAI research
Hugging Face daily papers·11d agoA Zeroth-Order Paradigm for LLM Preference Alignment#likelihood-displacement#post-training#preference-alignment1
arXiv cs.AI / cs.LG / cs.CL·12d agoInoculation Midtraining with Learned Neologisms#alignment#fine-tuning#llmAI safety & security
Hugging Face daily papers·13d agoNot All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training#grpo#math-reasoning#multimodal-llm
Hugging Face daily papers·13d agoRegister Tokens for Bounded-State Reasoning in Diffusion Language Models#diffusion-language-models#dream#llada1
Hugging Face daily papers·14d agoLightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition#capability-composition#dopd#efficient-reasoning
Hugging Face daily papers·16d agoExpert-Space Exploration in MoE Reinforcement Learning#expert-routing#grpo#mixture-of-experts
Hugging Face daily papers·16d agoMInTRL: Off-policy Intervention can boost On-policy RL#advantage-regression#exploration#on-policy
arXiv cs.AI / cs.LG / cs.CL·16d agoRetroThinker: Enabling Retrospective Thinking in Speech LLMs#chain-of-thought#dpo#gsm8kAI research1
arXiv cs.AI / cs.LG / cs.CL·18d agoActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation#benchmark#grpo#llmAI research1
Hugging Face daily papers·19d agoEliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation#alignment#distillation#post-training
Hugging Face daily papers·19d agoActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation#benchmark#grpo#llm
Hugging Face daily papers·19d agoMiles v0.1: Production-Level Post-Training#agentic-rl#distributed-training#open-source
Hugging Face daily papers·19d agoFeyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models#ctf#cyber-agents#cybergym
Hugging Face daily papers·20d agoRevisiting Complete Reasoning Traces for Post-Training#data-efficiency#distillation#llm
Hugging Face daily papers·21d agoMulti-Grid Post-Training for Long-Form Multi-Shot Video Generation#diffusion#long-video#multi-shot
Hugging Face trending models·21d agoTokenRhythm/NeoHorse-1-4B — new model trending #30 on Hugging Face#4b#agentic#llmModel release 15 min1
Hugging Face daily papers·23d agoOccamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work#35b#agentic-ai#co-work-agents
Hugging Face daily papers·25d agoVerify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation#grpo#llm#on-policy-distillation