Hugging Face daily papers·1d agoSparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference#llm-inference#pruning#sparsity