arXiv cs.AI / cs.LG / cs.CL·4d agoFlash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs#diffusion-llm#kv-cache#inferenceAI research 2 sources1