Semifactual Credit-Augmented Policy Optimization
SCAPO uses semifactual token stability to beat GRPO on AIME for two Qwen3 base models.
The paper studies how verifiable-reward reasoning models respond to semifactual prompt changes that preserve the problem and its answer. It finds token-level probability drift, and that suppressing high-drift candidates during decoding can improve accuracy without weight updates, exposing a weakness of GRPO's uniform token advantage. SCAPO uses normalized semifactual stability to reduce advantages for unstable tokens early in training without granting extra credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, it improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points and leads most math and all evaluated out-of-distribution benchmarks.
- Semifactual edits reveal token sensitivity to irrelevant prompt features.
- SCAPO reduces early-training credit for unstable response tokens.
- AIME accuracy rises 5.63 and 4.17 points over GRPO.
- Leads most math and all tested out-of-distribution benchmarks.
Full article223 words · extracted from arxiv.org · click to collapse
Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we introduce Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model scales, SCAPO achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods. These results suggest that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR. The code is available at https://github.com/DtYXs/SCAPO.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.40360