Alibaba's Qwen team launches Qwen3.8-Omni-Flash (1M-context omni-modal, API-only); ByteShape ships ShapeLearn GGUF quants for Qwen 3.8 27B
Alibaba's Qwen team announced Qwen3.8-Omni-Flash, an API-only omni-modal model with a 1M-token context window, agentic audio-video understanding, and tool use, claiming a 25%+ average improvement over Qwen3.5-Omni-Plus across 29 evaluations; no open weights…
Two developments hit the Qwen 3.8 family within roughly a day. Alibaba's Qwen team announced Qwen3.8-Omni-Flash (also written 'Qwen 3.8 Omni Flash') via its official blog, surfacing on Hacker News with 45 points and 7 comments. Follow-up details from MarkTechPost describe it as an API-only omni-modal model built on the Qwen3.8-Flash-Next architecture that accepts text, images, audio, and video and returns text, with a 1M-token context window and thinking enabled by default. It is positioned for fast, efficient multimodal inference: Qwen claims a 25%+ average improvement over Qwen3.5-Omni-Plus across 29 evaluations, with OmniVideoBench rising from 63.4 to 67.8 while using about 45.7% fewer tokens via coarse-to-fine agentic perception that gathers evidence instead of processing long video linearly. Qwen claims audio-visual performance near Gemini 3.8 Flash and audio performance above it. Audio input is supported in 113 languages, video up to 2 hours/2 GB, and availability spans 6 regions. The model is hosted on QwenCloud, Model Studio, and Qwen Studio at $0.15/$0.47 per 1M input/output tokens. No open weights were released, but Qwen open-sourced Qwen-MM-Plugins under Apache-2.0, adding omni-memory, omni-video2note, and omni-chatcut skills for agent harnesses. In parallel, ByteShape released its full ShapeLearn GGUF quantization set for Qwen 3.8 27B (base model released August 14, 2026), following earlier ShapeLearn-Lite quants published four days after launch. Five quants spanning IQ2_XXS 2.56bpw to IQ4_XS 3.84bpw were benchmarked on six GPUs against Unsloth Dynamic v3, ISTA-DASLab, Bartowski, and AtomicChat; the GPU-5 IQ4_XS quant reaches 99.63% of the BF16 aggregate score at roughly 90 tok/s on RTX Pro 6000 and RTX 5090, at 13.1 GB VRAM. Each GGUF bundles an MTP draft head, and a separate 1.1 GB DFlash2 draft model enables faster text-only speculative decoding via llama.cpp. The reports cover different models within the Qwen 3.8 family and do not conflict.
- Alibaba's Qwen team announced Qwen3.8-Omni-Flash via its official blog on 2026-09-17; it drew 45 points and 7 comments on Hacker News
- Qwen3.8-Omni-Flash is API-only, built on the Qwen3.8-Flash-Next architecture, accepts text/images/audio/video and returns text, with a 1M-token context window and thinking enabled by default
- Qwen claims 25%+ average improvement over Qwen3.5-Omni-Plus across 29 evaluations; OmniVideoBench rose from 63.4 to 67.8 with ~45.7% fewer tokens via coarse-to-fine agentic perception
- Qwen claims audio-visual performance near Gemini 3.8 Flash and audio performance above it
- Audio input in 113 languages; video up to 2 hours/2 GB; available in 6 regions
- Pricing: $0.15/$0.47 per 1M input/output tokens on QwenCloud, Model Studio, and Qwen Studio
- No open weights for the model; Qwen open-sourced Qwen-MM-Plugins under Apache-2.0 with omni-memory, omni-video2note, and omni-chatcut skills
- ByteShape released full ShapeLearn GGUF quants for Qwen 3.8 27B (base model released August 14, 2026), following ShapeLearn-Lite quants published four days after launch
Coverage timelineoldest first · each row is one article
- · 16h agoAlibaba releases Qwen 3.8 Omni Flash
Hacker News · AI· 55
Alibaba's Qwen team releases Qwen 3.8 Omni Flash, a fast omni-modal model announced on the official Qwen blog.
- · 13h agoShapelearn Qwen 3.8 27B (13.1 GB VRAM)
Hacker News · AI· 30
ByteShape released full ShapeLearn GGUF quants of Qwen 3.8 27B; its GPU-5 IQ4_XS reaches 99.63% of BF16 quality at 13.1 GB VRAM.
- · 6h agoAlibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use
MarkTechPost· 62
Alibaba's Qwen team launched Qwen3.8-Omni-Flash, an API-only omni-modal model with 1M-token context and agentic audio-video understanding and tool use.