Semifactual Credit-Augmented Policy Optimization
SCAPO adds semifactual token stability to GRPO, lifting Qwen3 AIME accuracy by up to 5.63 points.
SCAPO is a GRPO variant for reinforcement learning with verifiable rewards that uses semifactual prompt interventions to measure token-level probability drift. It lowers advantages for relatively unstable tokens early in training and does not award extra credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base it improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points. At both scales it is best on most math benchmarks and all reported out-of-distribution benchmarks among the compared methods.
- Semifactual edits preserve the problem and its answer
- GRPO assigns the same advantage to every response token
- SCAPO reduces credit for tokens that drift under those edits
- Qwen3-4B-Base gains 5.63 AIME points over GRPO
- Qwen3-1.7B-Base gains 4.17 points and leads most tests
Full article223 words · extracted from huggingface.co · click to collapse
Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates during decoding improves reasoning accuracy without updating model weights. These findings highlight a limitation of Group Relative Policy Optimization (GRPO), which assigns the same outcome-derived advantage to every response token and may reinforce potential spurious dependence alongside useful reasoning. Motivated by this observation, we introduce Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of GRPO that incorporates semifactual stability into token-level credit assignment. SCAPO measures token probability drift for fixed responses under semifactual interventions and uses normalized stability scores to reduce advantages for relatively unstable tokens during early training, while granting no additional credit for stability alone. On Qwen3-4B-Base and Qwen3-1.7B-Base, SCAPO improves AIME 2024-2026 accuracy over GRPO by 5.63 and 4.17 percentage points, respectively. At both model scales, SCAPO achieves the best results on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks among the compared methods. These results suggest that semifactual stability provides an effective training signal for improving reasoning and generalization through finer-grained credit assignment in RLVR. The code is available at https://github.com/DtYXs/SCAPO.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.40360