Disaggregated Quantization: Specializing LLM Prefill and Decode
Disaggregated quantization tailors LLM prefill and decode formats, raising accuracy and speeding time-to-first-token.
The paper proposes disaggregated quantization, specializing computation formats, weights, and storage separately for LLM prefill and decode. On Qwen 3 and Gemma 3, removing activation quantization during decode improved accuracy on decode-heavy tasks without increasing inference cost, while separate prefill weights matched or exceeded weight-only accuracy at 2-3-bit decode. An NVFP4 prefiller trained for released Qwen3.8-27B GGUF decoders improved 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro. Offloaded prefill from SSD delivered a 1.78x time-to-first-token speedup at 8K context in llama.cpp, with further tests in vLLM and on models up to 2.8T parameters.
- Disaggregated quantization uses different formats and weights for prefill and decode.
- Dropping decode activation quantization improved accuracy without added inference cost.
- An NVFP4 prefiller gained 32.5 MMLU-Pro and 35.3 MMMU-Pro points at 1-bit.
- SSD-offloaded prefill was 1.78x faster to first token at 8K in llama.cpp.
- Evaluation included vLLM serving and models up to 2.8 trillion parameters.
Full article180 words · extracted from huggingface.co · click to collapse
Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while matching or exceeding its accuracy at 2-3-bit decode on both decode-heavy and prefill-heavy tasks. With released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint. To accommodate the additional checkpoint on a single device, offloaded disaggregated prefill (ODP) streams its weights from SSD, amortizing loading over prompt length. On the same 27B model, ODP delivers a 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp. We evaluate accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.26333