arXiv cs.AI / cs.LG / cs.CL·4d agoTrain Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning#quantization#distillation#on-policyAI research
MarkTechPost·10d agoNunchux AI Introduces VC-Attention: A Training-Free Low-Bit Attention Kernel That Speeds Up Video Diffusion Transformers#attention-kernel#inference-optimization#low-bit 5 min1