Hugging Face daily papers·3d agoAV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation#diffusion#reinforcement-learning#audio-video-generation
Hugging Face daily papers·3d agoWanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation#text-to-video#prompt-enhancement#wanpe
arXiv cs.AI / cs.LG / cs.CL·3d agoWhen and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment#grpo#on-policy-distillation#reinforcement-learningAI research
arXiv cs.AI / cs.LG / cs.CL·4d agoSpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue#conversational-memory#multi-party-dialogue#reinforcement-learningAI research 2 sources
Hugging Face trending models·4d agoXiaomiMiMo/MiMo-V2.6-Flash-RL — new model trending #30 on Hugging Face#xiaomi#mimo-v2.6#mixture-of-expertsModel release 5 sources 7 min1
arXiv cs.AI / cs.LG / cs.CL·8d agoCodeMidas: Scaling Agentic Coding RL Environments from Code Itself#reinforcement-learning#coding-agents#rl-environmentsAI research1
arXiv cs.AI / cs.LG / cs.CL·8d ago$λ$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource#grpo#flow-matching#reinforcement-learningAI research
Hugging Face daily papers·9d agoCodeMidas: Scaling Agentic Coding RL Environments from Code Itself#code-generation#codemidas#coding-agents2
arXiv cs.AI / cs.LG / cs.CL·9d agoDon't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL#grpo#llm-agents#qwen3AI research
arXiv cs.AI / cs.LG / cs.CL·9d agoMulti-Dimensional Prosody Judgment For Live Streaming Speech Synthesis#distillation#grpo#llm-judgeAI research1
MarkTechPost·10d agoGoogle Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out#diffusion-models#embeddings#google-research 4 min1
Hugging Face daily papers·10d agoDon't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL#actobs#exploration#grpo
Hugging Face daily papers·13d agoNot All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training#grpo#math-reasoning#multimodal-llm
arXiv cs.AI / cs.LG / cs.CL·15d agoExpert-Space Exploration in MoE Reinforcement Learning#expert-routing#grpo#llm-trainingAI research
Hugging Face daily papers·16d agoLearning to Solve Hard Problems in RL for LLMs by Never Giving Up#adaptive-sampling#coding-benchmark#grpo
Hugging Face daily papers·16d agoExpert-Space Exploration in MoE Reinforcement Learning#expert-routing#grpo#mixture-of-experts
Hugging Face Blog·17d agoAsync GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL#distributed-training#fine-tuning#grpoAI tools & infra
arXiv cs.AI / cs.LG / cs.CL·18d agoActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation#benchmark#grpo#llmAI research1
arXiv cs.AI / cs.LG / cs.CL·18d agoThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR#grpo#prompt-selection#qwen2.5-mathAI research1
Hugging Face daily papers·19d agoActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation#benchmark#grpo#llm
Hugging Face daily papers·22d agoDataFlex-RL: An Evaluation Platform for RLVR Data Policies#data-policy#evaluation#grpo
arXiv cs.AI / cs.LG / cs.CL·22d agoMulti-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe#benchmark#data-synthesis#fine-tuningAI research1
Hugging Face Blog·24d agoFine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps#fine-tuning#grpo#hugging-faceAI tools & infra
Hugging Face daily papers·25d agoVerify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation#grpo#llm#on-policy-distillation