Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
WeightPE finetunes int8 weights so Re-Pair grammars shrink to 0.38–0.43x of QAT, costing about 1–2 accuracy points.
Weight Pair Encoding finetunes int8 network weights so a lossy Re-Pair compressor finds a smaller grammar, using a straight-through estimator under a global L2 budget. On MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, grammars are 0.43x and 0.38x the size of an equivalent int8 QAT run, costing 1.9 and 1.1 accuracy points. The same trend appears with LZ78 and SEQUITUR, compressors the networks were not trained against. The authors call grammar size a new explicit training objective for weights.
34