Latent Space·3d ago🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)#ai-arms-race#bio-security#dual-use 4 min
arXiv cs.AI / cs.LG / cs.CL·4d agoThe Sirens' Song: When Proximal Background Context Overshadows Distant Evidence#long-context#llm#lyraAI research
arXiv cs.AI / cs.LG / cs.CL·5d agoDolphinBench: Mapping the Pareto Frontier of Agent Memory#dolphinbench#agent-memory#benchmarkAI research
MarkTechPost·8d agoAlibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use#agentic-ai#alibaba#long-context 5 min1
arXiv cs.AI / cs.LG / cs.CL·9d agoOn-Demand Attention: Language Models Know When to Recall#attention#decoding#inference-optimizationAI research
Hacker News · AI·10d agoDeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression#deepseek#deepseek-v4.1-flash#fp4Model release 15 min1
Hugging Face daily papers·10d agoDeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression#deepseek#inference-efficiency#kv-cache1
Hugging Face trending models·11d agoXingChen-AGI/Xing4.0-29B-A4B — new model trending #30 on Hugging Face#ascend-npu#china-telecom#long-contextModel release 7 min1
Hugging Face daily papers·12d agoFathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches#decoding#host-memory-offload#inference-optimization1
Hugging Face daily papers·14d agoSpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization#continual-pretraining#gated-deltanet#linear-attention2
Hugging Face daily papers·14d agoFlattening Every Memory Peak in Long-Context Mixture-of-Experts Training#distributed-training#long-context#memory-optimization1
Hugging Face trending models·14d agoyandex/AliceAI-Foundation-80B-A3B-Base — new model trending #30 on Hugging Face#aliceai#base-model#llmModel release 15 min
Latent Space·15d ago[AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale#deepseek#encoder-decoder#inference-efficiency 15 min1
arXiv cs.AI / cs.LG / cs.CL·15d agoSAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking#attention-sparsification#efficiency#flashattentionAI research2
Hugging Face trending models·15d agoAgnes-AI/Agnes-3.0-Flash — new model trending #30 on Hugging Face#agnes#apache-2.0#llmModel release 12 min
Hugging Face daily papers·16d agoSAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking#attention#efficiency#long-context1
Hugging Face daily papers·16d agoZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search#7b#agentic#foundation-model1
MarkTechPost·17d agoDeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse#deepseek#deepseek-v4.1-flash#kv-cache 4 min2
arXiv cs.AI / cs.LG / cs.CL·17d agoConvMem: Convolutional Memory for Long-Context Reasoning#hierarchical-convolution#inference#llmAI research
arXiv cs.AI / cs.LG / cs.CL·18d agoLearning Length-Extrapolatable Recurrent Models#bptt#cst#length-extrapolationAI research1
arXiv cs.AI / cs.LG / cs.CL·19d agoKalman Delta Networks: Uncertainty-aware Associative Memory#architecture#associative-memory#kalman-filterAI research1
Hugging Face daily papers·21d agoPARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents#llm-agents#long-context#multi-hop-qa1
Hugging Face trending models·25d agoIFM/K2-Horizon-MoVA-36B-A4B — new model trending #15 on Hugging Face#512k-context#hugging-face#ifmModel release 15 min
Hugging Face trending models·Aug 26, 2026unsloth/Qwen3.8-Flash-Next-GGUF — new model trending #21 on Hugging Face#gguf#long-context#moeModel release 15 min1