Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Weights
WeightPE finetunes int8 weights so Re-Pair grammars shrink to 0.38–0.43x of QAT, costing about 1–2 accuracy points.
Weight Pair Encoding finetunes int8 network weights so a lossy Re-Pair compressor finds a smaller grammar, using a straight-through estimator under a global L2 budget. On MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, grammars are 0.43x and 0.38x the size of an equivalent int8 QAT run, costing 1.9 and 1.1 accuracy points. The same trend appears with LZ78 and SEQUITUR, compressors the networks were not trained against. The authors call grammar size a new explicit training objective for weights.
- Straight-through estimator trains through lossy Re-Pair rewritten weights
- ViT-B/16 grammar is 0.43x QAT size, minus 1.9 accuracy points
- ViT-L/16 grammar is 0.38x QAT size, minus 1.1 accuracy points
- Gains also appear under LZ78 and SEQUITUR without matched finetuning
Full article168 words · extracted from arxiv.org · click to collapse
We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and 1.1 accuracy points. The trend extends to different grammar compressors (LZ78, SEQUITUR), over which the networks has not be finetuned against. To our knowledge, this is the first time grammar size has been used as an explicit training objective for network weights.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.31564