TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning
TRIAGE stabilizes native NVFP4 RL on Qwen3 models, matching full precision with up to 2.3x rollout throughput.
TRIAGE addresses instability when reinforcement learning for large language models runs in native NVFP4, where learner-sampler mismatch can skew policy-gradient updates. The authors find an early imbalance favoring negative-advantage, negative-gap updates, with tail tokens concentrating in a few response segments before mismatch spreads. TRIAGE uses segment-level diagnosis to rebalance updates and applies bounded repair while keeping weight-and-activation 4-bit forward passes on both sampler and learner. On Qwen3-4B and Qwen3-30B-A3B it stays stable, matches full-precision results on five math reasoning benchmarks, and delivers up to 2.3x higher rollout throughput than BF16.