ZeroHour
Story · 3 sources · 3 articlesfirst updated ()

DeepSeek releases open-weight V4.1-Flash: 1M-token context, FP4 KV cache, and novel causal encoder-decoder MoE

infoModel releaseimportance 80
What's new: First merged story for this dashboard (no previous summary) — all items are new coverage of the DeepSeek-V4.1-Flash release on or around 2026-09-10.|Memory efficiency delta vs prior state: KV cache cut to about 1/4 of DeepSeek-V4-Flash in GPU memory (up to 1/8 offloaded) and 437x smaller per token than DeepSeek-V1, via FP4 (E2M1) quantization and the new causal encoder-decoder/CSA2 design.|New…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

DeepSeek released DeepSeek-V4.1-Flash, an open-weight MIT-licensed multimodal MoE with a 1M-token context and FP4 KV cache that cuts memory to roughly a quarter of its predecessor; reports describe it as a 552B backbone plus 196B Engram parameters (one report…

On or around September 10, 2026, DeepSeek released DeepSeek-V4.1-Flash as open weights under the MIT license on Hugging Face, with vLLM and SGLang support, day-0 serving from Baseten, an Ollama rollout to paid subscribers, and API availability at V4-Flash prices (one report cites $0.30 per 1M input tokens, $1.20 per 1M output tokens, and a 50% off-peak discount). The reports disagree on total size: MarkTechPost and The Decoder describe a 552B-parameter MoE backbone, with MarkTechPost adding 196B 'Engram' parameters, while Latent Space describes 763B total parameters; all agree it activates 8B parameters at prefill and 16B at decode. The model uses a novel causal encoder-decoder design with Compressed Sparse Attention 2 (CSA2), which shares KV across layers in Full, Reindex, and Reuse modes, and stores the KV cache in FP4 (E2M1), cutting it to 890 bytes per token — about 1/4 of DeepSeek-V4-Flash (characterized as ~1/4 in GPU memory and up to 1/8 when offloaded) and 437x smaller than DeepSeek-V1. It supports a 1M-token context with text and image input and was pre-trained from scratch on 45 trillion tokens. Benchmarks: 90.6 on Terminal-Bench 2.1 and 74.2% on DeepSWE v1.1, ahead of Opus 5 and GPT-5.6 Sol; a 40 score on the Artificial Analysis Intelligence Index, above DeepSeek V4 Pro 0813; and the #1 open-weight rank on the Vals Index, ahead of Kimi K3. Caveats from The Decoder: the gains are attributed to data and RL scaling rather than new algorithms, and gaps remain versus closed frontier models on expert tasks and complex image understanding.

  • DeepSeek-V4.1-Flash released on or around 2026-09-10 as open weights under the MIT license on Hugging Face, with vLLM and SGLang support; also available via API at V4-Flash prices.
  • Size reported inconsistently: 552B-parameter backbone plus 196B Engram parameters (MarkTechPost; The Decoder cites 552B total) vs 763B total parameters (Latent Space); all reports agree on 8B active parameters at prefill and 16B at decode.
  • Novel causal encoder-decoder architecture; Compressed Sparse Attention 2 (CSA2) shares KV across layers in Full, Reindex, and Reuse modes.
  • FP4 (E2M1) KV cache quantization to 890 bytes per token — about 1/4 of DeepSeek-V4-Flash (reported as ~1/4 in GPU memory and up to 1/8 offloaded) and 437x smaller than DeepSeek-V1.
  • 1M-token context window with text and image input; pre-trained from scratch on 45 trillion tokens.
  • Benchmarks: 90.6 on Terminal-Bench 2.1; 74.2% on DeepSWE v1.1, ahead of Anthropic Opus 5 and OpenAI GPT-5.6 Sol; 40 on the Artificial Analysis Intelligence Index, above DeepSeek V4 Pro 0813; #1 open-weight model on the Vals Index, ahead of…
  • API pricing per one report: $0.30 per 1M input tokens, $1.20 per 1M output tokens, with a 50% off-peak discount; day-0 support from Baseten and Ollama rollout to paid subscribers.
  • Caveats: gains attributed to data and RL scaling rather than new algorithms; gaps remain versus closed frontier models on expert tasks and complex image understanding (The Decoder).

Coverage timeline

  1. · 6d ago
    MarkTechPost· 80
    DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse

    DeepSeek released open-weight V4.1-Flash, a 552B MoE model with 1M context and FP4 KV cache, beating Opus-5 and GPT-5.6 Sol on agent benchmarks.

  2. · 6d ago
    The Decoder· 76
    New Deepseek model V4.1-Flash cuts memory needs for AI agents

    DeepSeek released V4.1-Flash, a 552B-parameter open-weight model cutting KV cache needs to a quarter of its predecessor for cheaper million-token AI agents.

  3. · 4d ago
    Latent Space· 80
    [AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale

    DeepSeek released V4.1-Flash, an open-weight 763B-parameter model with a novel causal encoder-decoder architecture, 1M context, vision input, and MIT license.