arXiv cs.AI / cs.LG / cs.CL·12d agoBellman Policy Optimization#llm-reasoning#math-reasoning#policy-optimizationAI research
Hugging Face daily papers·13d agoNot All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training#grpo#math-reasoning#multimodal-llm
Hugging Face daily papers·16d agoLearning to Solve Hard Problems in RL for LLMs by Never Giving Up#adaptive-sampling#coding-benchmark#grpo
Hugging Face daily papers·16d agoZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search#7b#agentic#foundation-model1
Hugging Face daily papers·24d agoFlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience#math-reasoning#qwen3#reasoning-models