GGUF vs GPTQ vs AWQ vs EXL2: LLM Model Formats Explained (2026)
An explainer separates LLM storage containers (safetensors, GGUF) from quantization methods (GPTQ, AWQ, EXL2/EXL3), with memory math and quality trade-offs.
The article explains that containers like safetensors, GGUF, and PyTorch pickle define on-disk tensor storage, while GPTQ, AWQ, bitsandbytes NF4, and llama.cpp K/I-quants define quantization methods, with EXL2/EXL3 combining both. GGUF, introduced August 2023 by Georgi Gerganov, replaced GGML using typed key-value metadata and can embed tokenizers and chat templates; a Q4_K_M quant averages about 4.5 bits per weight. GPTQ, published at ICLR 2023 by Frantar, Ashkboos, Hoefler, and Alistarh, quantized a 175B-parameter model in roughly 4 GPU hours to 3-4 bits with ~3.25x speedup on A100 GPUs. It also notes safetensors eliminates pickle's arbitrary code execution risk when loading untrusted checkpoints.