Learning Functional Subspaces for Neural Network Compression
Learnable Subspace Projections compress LLMs end-to-end, beating baselines at 70% compression on Llama-2-7B.
Learnable Subspace Projections (LSP) compress transformers by learning which weight subspaces to discard, instead of using local closed-form criteria that ignore how errors propagate with depth. Orthogonal projectors are optimized jointly against output KL or the original loss while pretrained weights stay frozen, then merged into standard low-rank factors. Across OPT-125M/1.3B, Qwen3-4B, Llama-2-7B, and ViT-B/16, LSP's advantage grows with compression. At 70% compression, Llama-2-7B reaches 10.9 WikiText-2 perplexity and 42.2% mean zero-shot accuracy versus 13.3 and 36.0% for the strongest baseline, with up to 1.6x faster decoding and 13.5x lower combined weight and KV-cache memory at a 128k context.