Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features
SCALE gates frozen SFT features with entropy-trained controls, improving math and code scores on three Qwen models.
The paper argues that existing SFT token-reweighting methods use nonnegative coefficients, so they can suppress or amplify updates but cannot reverse harmful features or extrapolate a fixed SFT direction. SCALE freezes the pretrained model and the SFT delta, then learns bounded token- and module-specific gates by minimizing predictive entropy alone. On Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Qwen3-4B-Base it reports math averages of 37.84, 43.60, and 36.57, above the strongest baselines, and the best average code-generation scores on HumanEval, HumanEval+, and MBPP.