arXiv cs.AI / cs.LG / cs.CL·2d agoAD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control#world-models#model-predictive-control#roboticsAI research
arXiv cs.CR·2d agoForgetMimic: Motion Unlearning for Reinforcement Learning Humanoid Control#forgetmimic#machine-unlearning#reinforcement-learningAI research 2 sources
arXiv cs.AI / cs.LG / cs.CL·2d agoPoEM: Predicting RL Outcomes from Existing Policies#poem#reinforcement-learning#reward-modelsAI research
Hugging Face daily papers·2d agoJust Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures#ai-safety#alignment#jailbreak 2 sources
Hugging Face daily papers·3d agoAV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation#diffusion#reinforcement-learning#audio-video-generation
Hugging Face daily papers·3d agoRufus-Air: An Open LLM Post-Training Recipe#post-training#sft#reinforcement-learning
Hugging Face daily papers·3d agoIterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis#llm-agents#deep-search#reinforcement-learning
Hugging Face daily papers·3d agoQwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents#qwen#mobile-agents#reinforcement-learning
arXiv cs.AI / cs.LG / cs.CL·3d agoWhen and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment#grpo#on-policy-distillation#reinforcement-learningAI research
arXiv cs.AI / cs.LG / cs.CL·3d agoLEAP-CBF: A Safety Filter for Uncertain Systems with Least-Effort Adversarial Potentials#control-barrier-functions#reinforcement-learning#roboticsAI research
The Decoder·3d agoAnthropic engineer explains why Claude's writing got worse although the model got smarter#anthropic#claude#opus 2 min
MarkTechPost·4d agoKyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning#kyutai#voice-of-reason#glm-4-voice 4 min
TechCrunch · AI·4d agoSnorkel AI triples valuation to $3.5B as demand for AI training data booms#ai#data#reinforcement-learning 2 min
Hugging Face daily papers·4d agoVerifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms#agentic-rl#environment-generation#qwen
arXiv cs.AI / cs.LG / cs.CL·4d agoA Decentralized Partially Observable Team Decision Methodology with Delayed Information Sharing#multi-agent#pomdp#reinforcement-learningAI research
arXiv cs.AI / cs.LG / cs.CL·4d agoSpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue#conversational-memory#multi-party-dialogue#reinforcement-learningAI research 2 sources
arXiv cs.AI / cs.LG / cs.CL·4d agoOptimal Sequential Annotations for Off-Policy Evaluation#off-policy-evaluation#reinforcement-learning#annotationAI research
arXiv cs.AI / cs.LG / cs.CL·4d agoBeyond Repeated Sampling: Learning Search Policies for LLM Reasoning#llm#reasoning#test-time-computeAI research
arXiv cs.AI / cs.LG / cs.CL·4d agoMAGIC: Mixed-Granularity Agent Graphs via Incremental Construction with Dense-Reward Reinforcement Learning#multi-agent#reinforcement-learning#llmAI research
Latent Space·4d ago[AINews] Xiaomi MiMo-V2.6-Pro 1T-A42B: the new top Open Weights model, trained for $3M#xiaomi#mimo#open-weights 5 sources 13 min
Hugging Face daily papers·5d agoPACT: From Credit Assignment to Critic Alignment#reinforcement-learning#llm-post-training#actor-critic
arXiv cs.AI / cs.LG / cs.CL·5d agoCritical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use#reinforcement-learning#tool-use#multi-turnAI research
arXiv cs.AI / cs.LG / cs.CL·5d agoVisuomotor Robotic Pruning in Planar Orchards Using Hybrid Reinforcement Learning#reinforcement-learning#robotics#visuomotorAI research
arXiv cs.CR·5d agoReinforcement Learning Inspired Black-box Adversarial Attacks for Computer Vision#adversarial-attack#black-box#reinforcement-learningAI safety & security