Semifactual Credit-Augmented Policy Optimization
SCAPO uses semifactual token stability to beat GRPO on AIME for two Qwen3 base models.
The paper studies how verifiable-reward reasoning models respond to semifactual prompt changes that preserve the problem and its answer. It finds token-level probability drift, and that suppressing high-drift candidates during decoding can improve accuracy without weight updates, exposing a weakness of GRPO's uniform token advantage. SCAPO uses normalized semifactual stability to reduce advantages for unstable tokens early in training without granting extra credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, it improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points and leads most math and all evaluated out-of-distribution benchmarks.