Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization
ZIP-SR and ZE-EDEN cut 4-bit AdamW's validation-loss gap to 32-bit AdamW by up to 70%.
The paper redesigns 4-bit AdamW optimizer-state quantization around the coordinate used for rounding. ZIP-SR keeps zero in the second-moment codebook and computes stochastic-rounding probabilities in preconditioner space, while ZE-EDEN uses a zero-excluding codebook and rescales quantized blocks. Both use 4-bit NormalFloat for the first moment. Across GPT- and Llama-style pretraining from 130M to 2.7B parameters, both reduce TorchAO's mean validation-loss gap to 32-bit AdamW at every size, with the largest reduction reaching 70%, and they remain close to full precision in supervised fine-tuning.
- Quantization error in Adam moments distorts later preconditioners.
- ZIP-SR rounds in preconditioner space and retains a zero code.
- ZE-EDEN rescales a zero-excluding second-moment codebook.
- Gap to 32-bit AdamW shrinks by up to 70% from 130M to 2.7B.
Full article225 words · extracted from arxiv.org · click to collapse
Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates. We redesign 4-bit optimizer-state quantization for AdamW from the perspective of \emph{rounding space}: the coordinate in which a quantizer chooses between adjacent reconstruction levels. For the second moment, a local analysis of the quantization cell adjacent to zero shows that small mean state error need not imply small mean preconditioner error at the next step. A one-dimensional quadratic construction further shows qualitatively different optimization dynamics under state-space and preconditioner-space rounding. These results motivate Zero-Inclusive Preconditioner-space Stochastic Rounding (\textbf{ZIP-SR}), which retains zero in the second-moment codebook and computes stochastic-rounding probabilities in preconditioner space. As a complementary route, Zero-Excluding EDEN calibration (\textbf{ZE-EDEN}) uses a zero-excluding second-moment codebook and rescales the quantized second-moment block to mitigate the preconditioner distortion caused by the positive quantization floor. Both configurations use 4-bit NormalFloat (NF4) for the first moment, with targeted stochastic rounding of the LM-head first moment during the final 10\% of training. Across GPT- and Llama-style pretraining experiments ranging from \textbf{130M} to \textbf{2.7B} parameters, both methods reduce TorchAO 4-bit AdamW's mean validation-loss gap to 32-bit AdamW at every evaluated model size, with the largest reported gap reduction reaching \textbf{70\%}. In full-parameter supervised fine-tuning, both recipes achieve lower validation loss than TorchAO while remaining close to 32-bit AdamW on downstream tasks.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.12444