Open-weight labs ship new agentic models: Nex-AGI's trillion-parameter Nex-N2.5 family and DeepSeek's KV-cache-lean V4.1-Flash
Two separate open-weight releases in one week target long-horizon AI agents: Nex-AGI's Nex-N2.5 family (Max on a 1.6T-parameter MoE) and DeepSeek's 552B-parameter V4.1-Flash with 1M-token context and a KV cache cut to about a quarter of its predecessor.
On 2026-09-08, Nex-AGI announced Nex-N2.5, a next-generation family of agentic models in three sizes (mini, Pro, Max) focused on long-horizon tasks including computer use, web browsing, and autonomous program execution. Nex-N2.5-Max is built on a 1.6-trillion-parameter text-only Mixture-of-Experts foundation, the company's first complete post-training effort at trillion-parameter scale. Weights were listed as "coming soon" at publication, promised open-source on Hugging Face and ModelScope, with hosted access via OpenRouter. Nex-N2.5-Pro scored 82.7 on Terminal-Bench 2.1 and 61.2 on SWE-Bench Pro in comparisons against Claude Opus 5, GPT-5.6 Sol, Kimi-K3, GLM-5.3, DeepSeek-V4-Pro-0813, and Qwen3.8-Max. Two days later, on 2026-09-10, DeepSeek released V4.1-Flash, a distinct release from a different company: a 552-billion-parameter multimodal model with a 1 million-token context, trained from scratch on 45 trillion tokens of text and images. V4.1-Flash cuts KV cache footprint to roughly a quarter of DeepSeek-V4-Flash in fast GPU memory and one-eighth when offloaded — 437x smaller per token than DeepSeek-V1 — via an encoder/decoder split, 8-16B active parameters per token, and FP4 cache storage. It scores 74.2% on DeepSWE v1.1, narrowly beating Anthropic Opus 5 and OpenAI GPT-5.6 Sol, with gains attributed to data and RL scaling rather than new algorithms. Weights are on Hugging Face under MIT license, also served via API at V4-Flash prices. The two reports describe separate launches and do not contradict each other; both position open-weight models as competitive with closed frontier models on agentic benchmarks, while DeepSeek notes remaining gaps versus closed frontier models on expert tasks and complex image understanding.
- Nex-AGI's Nex-N2.5 family (announced 2026-09-08) ships in three sizes: mini, Pro, and Max, focused on agentic computer use, web browsing, and autonomous program execution
- Nex-N2.5-Max is built on a 1.6-trillion-parameter text-only MoE foundation — Nex-AGI's first complete post-training effort at trillion-parameter scale
- Nex-N2.5-Pro scores 82.7 on Terminal-Bench 2.1 and 61.2 on SWE-Bench Pro, benchmarked against Claude Opus 5, GPT-5.6 Sol, Kimi-K3, GLM-5.3, DeepSeek-V4-Pro-0813, and Qwen3.8-Max
- Nex-N2.5 weights were listed as "coming soon" at publication; open weights promised on Hugging Face and ModelScope, with hosted access via OpenRouter
- DeepSeek V4.1-Flash (released 2026-09-10) is a 552B-parameter multimodal model with 1M-token context, trained from scratch on 45T tokens of text and images
- V4.1-Flash reduces KV cache to ~1/4 of DeepSeek-V4-Flash in fast GPU memory and ~1/8 when offloaded — 437x smaller per token than DeepSeek-V1 — via an encoder/decoder split, 8-16B active parameters per token, and FP4 cache storage
- V4.1-Flash scores 74.2% on DeepSWE v1.1, narrowly beating Anthropic Opus 5 and OpenAI GPT-5.6 Sol, with gains attributed to data and RL scaling rather than new algorithms
- V4.1-Flash is released under MIT license on Hugging Face and served via API at V4-Flash prices; DeepSeek notes remaining gaps vs. closed frontier models on expert tasks and complex image understanding
Coverage timelineoldest first · each row is one article
- · 7d agonex-agi/Nex-N2.5-Pro — new model trending #30 on Hugging Face
Hugging Face trending models· 62
Nex-AGI launches Nex-N2.5 agentic model family (mini/Pro/Max), with Max built on a 1.6-trillion-parameter MoE foundation.