Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces
Bias-only label-free test-time RL lifts Qwen2.5-7B to 76.67% on MATH-500 using about 100K parameters.
The paper introduces label-free bias-only test-time reinforcement learning that uses majority-vote pseudolabels and updates roughly 100,000 bias parameters while freezing the pretrained backbone. On MATH-500, Qwen2.5-7B reaches 76.67 percent accuracy, slightly above the authors' labeled bias-steering reproduction and using 76,000 times fewer parameters than full-parameter TTRL. The same procedure improves MathVista, AI2D, LogicVista, and MMAU, and learned steering vectors transfer to 4,500 held-out MATH problems. Analyses link gains to rollout consensus and gradient energy in the bias subspace.
- Only about 100,000 bias parameters are updated; the backbone stays frozen
- Qwen2.5-7B reaches 76.67 percent accuracy on MATH-500
- Uses 76,000 times fewer parameters than full-parameter TTRL
- Steering transfers to 4,500 held-out MATH problems and other modalities
Full article196 words · extracted from huggingface.co · click to collapse
Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the optimization space are severely restricted? We answer this question with label-free bias-only TTRL, which uses majority-vote pseudolabels as rewards and optimizes only ~100K bias parameters while keeping the pretrained backbone frozen. On MATH-500, our approach reaches 76.67% accuracy with Qwen2.5-7B, slightly exceeding our own labeled bias-steering reproduction while optimizing 76,000x fewer parameters than full-parameter TTRL. The same training procedure improves performance across vision-language and audio reasoning tasks, including MathVista, AI2D, LogicVista, and MMAU. We further show that the learned steering vectors transfer to 4,500 held-out MATH problems, indicating that the adaptation is not limited to the problems used during test-time optimization. Finally, we analyze why this highly restricted adaptation can work, showing that majority-vote reliability improves with rollout consensus and that bias subspaces with greater accessible gradient energy exhibit stronger downstream trainability. These results demonstrate that substantial test-time adaptation can emerge from optimizing a tiny bias-only subspace using entirely label-free rewards.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.18587