Softmax Reparameterization for Output-Head Quantization
A post-training softmax reparameterization cuts output-head quantization error, dropping Phi-4-mini AW-MSE KL from 0.936 to 0.256.
The paper proposes softmax reparameterization, a post-training method that selects a functionally equivalent language-model output head before quantization by subtracting a scaled vocabulary-row mean. The coefficient is chosen by validation KL separately for RTN, activation-weighted MSE, and GPTQ, preserving the full-precision softmax while leaving the decoder unchanged. On Phi-4-mini, W4 activation-weighted MSE KL falls from 0.936 to 0.256, and frozen WikiText coefficients transfer to C4 and OpenWebMath. With the decoder in BF16, quantizing the Phi output head reduces batch-one generation latency by 10.8%.
- One-dimensional search includes the original head and fixed mean-centering.
- Phi-4-mini W4 AW-MSE KL drops from 0.936 to 0.256.
- WikiText-selected coefficients transfer to C4 and OpenWebMath.
- Quantizing the Phi head cuts batch-one latency 10.8% versus BF16.
Full article232 words · extracted from huggingface.co · click to collapse
Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ. This one-dimensional search includes the original head and fixed mean-centering, preserves the full-precision softmax distribution, and leaves the trained decoder unchanged; a rank-one correction handles nonlinear logit paths such as soft-capping. Across seven heads, W4 gains concentrate where baseline quantization substantially distorts predictions: on Phi-4-mini, AW-MSE KL falls from 0.936 to 0.256. The gains survive stronger GPTQ calibration and remain complementary to exact per-channel scaling and affine quantization. Across four heads and three W4 quantizers, frozen WikiText-selected coefficients also transfer to C4 and OpenWebMath, outperforming mean-centering in all 18 comparisons where the frozen coefficient differs from 1 and matching it in the remaining six. At W2, used as a compression stress test, benefits broaden across nearly the full model--quantizer matrix. Matched residual analysis shows that improved fidelity can accompany greater logit reconstruction error while reducing the residual's Fisher-weighted cost. For shift-compatible heads, reparameterization adds no inference operation and preserves packed W4 execution: with the decoder held in BF16, quantizing the Phi output head reduces batch-one generation latency by 10.8% relative to the BF16-head baseline.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.31291