EasyPPO: Stabilizing the Critic Is Key
EasyPPO stabilizes the PPO critic for LLM training and outperforms vanilla PPO on coding, math, and search.
EasyPPO targets two critic failure modes that destabilize PPO when training large language models: filtering truncated rollouts from both actor and critic, and high-variance prompts dominating critic updates. It trains the critic on completed and truncated returns, weights each prompt's critic loss by the inverse standard deviation of sampled returns, and uses moderately smaller critic mini-batches. On FrontierCS coding, AIME24 math, and Search-R1 multi-turn search, it stays stable for the full horizon and beats vanilla PPO, VAPO, and HL-Gauss PPO. Best validation scores improve 14.89%, 2.28%, and 9.47% over PPO respectively.
- Critic instability is identified as a major PPO failure mode for LLMs.
- Actor-only overlong filtering keeps truncated returns in critic training.
- Noise-normalized regression balances high-variance prompts in critic updates.
- Relative gains versus PPO are 14.89%, 2.28%, and 9.47% on three tasks.
- Stable across FrontierCS coding, AIME24 math, and Search-R1 multi-turn search.
Full article203 words · extracted from huggingface.co · click to collapse
A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for large language models (LLMs). We identify two critic failure modes that destabilize PPO. First, filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, allowing truncation to increase even as conditional reward improves. Second, heterogeneous return noise can cause high-variance prompts to dominate critic updates in finite batches. We introduce EasyPPO to address these failures. Actor-only overlong filtering trains the critic on returns from both completed and truncated rollouts. Noise-normalized critic regression weights each prompt's critic loss by the inverse standard deviation of its sampled returns, balancing noise contributions across prompts. Moderately smaller critic mini-batches confine outlier influence to fewer rollouts during gradient clipping. Across continuous-reward coding on FrontierCS, binary-reward mathematical reasoning on AIME24, and multi-turn search on Search-R1, EasyPPO remains stable throughout the full training horizon and consistently outperforms vanilla PPO, VAPO, and HL-Gauss PPO. Its best validation scores show relative gains of 14.89%, 2.28%, and 9.47% over PPO, respectively.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.36802