arXiv cs.AI / cs.LG / cs.CL·4d agoFlash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs#diffusion-llm#kv-cache#inferenceAI research 2 sources1
Hacker News · AI·8d agoCache-to-Cache: Direct Semantic Communication Between LLMs (2025)#arxiv#inference#kv-cacheAI research1
arXiv cs.AI / cs.LG / cs.CL·9d agoOn-Demand Attention: Language Models Know When to Recall#attention#decoding#inference-optimizationAI research
Hacker News · AI·10d agoDeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression#deepseek#deepseek-v4.1-flash#fp4Model release 15 min1
Hugging Face daily papers·10d agoDeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression#deepseek#inference-efficiency#kv-cache1
arXiv cs.AI / cs.LG / cs.CL·10d agoLong-Lived Characters, Local Inference: Incremental Memory Maintenance for Game NPCs#agent-memory#game-npcs#kv-cacheAI research
arXiv cs.AI / cs.LG / cs.CL·11d agoJustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management#inference#kv-cache#llm-servingAI research
Hugging Face daily papers·12d agoFathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches#decoding#host-memory-offload#inference-optimization1
The Decoder·16d agoNew Deepseek model V4.1-Flash cuts memory needs for AI agents#ai-agents#deepseek#inference-efficiency 4 min1
MarkTechPost·17d agoDeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse#deepseek#deepseek-v4.1-flash#kv-cache 4 min2
Hugging Face daily papers·18d agoTRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents#gui-agents#inference-efficiency#kv-cache
Hugging Face daily papers·19d agoGrouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction#attention#gqa#inference-efficiency
Hugging Face daily papers·20d agoLatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay#gated-deltanet#hybrid-llm#inference
arXiv cs.CR·23d agoForgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents#agent-security#kv-cache#llm-agentsAI safety & security1
Hugging Face daily papers·23d agoBeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference#chain-of-thought#inference-optimization#kv-cache1
Hugging Face daily papers·25d agoShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding#inference-efficiency#kv-cache#mllm
Hugging Face daily papers·Aug 27, 2026HyQuant: Hybrid-Precision Quantization for LLM Attention#attention#efficiency#kv-cache1