Hacker News · AI·10d agoDeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression#deepseek#deepseek-v4.1-flash#fp4Model release 15 min1
Hugging Face daily papers·12d agoFathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches#decoding#host-memory-offload#inference-optimization1
arXiv cs.AI / cs.LG / cs.CL·15d agoSAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking#attention-sparsification#efficiency#flashattentionAI research2
MarkTechPost·17d agoDeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse#deepseek#deepseek-v4.1-flash#kv-cache 4 min2
Hugging Face trending models·17d agodeepseek-ai/DeepSeek-V4.1-Flash — new model trending #28 on Hugging Face#deepseek#deepseek-v4.1-flash#kv-cache-compressionModel release 10 min1
Hugging Face trending models·Aug 26, 2026unsloth/Qwen3.8-Flash-Next-GGUF — new model trending #21 on Hugging Face#gguf#long-context#moeModel release 15 min1