ZeroHour
Story · 2 sources · 3 articlesfirst updated ()

Alibaba's Qwen team launches Qwen3.8-Omni-Flash (1M-context omni-modal, API-only); ByteShape ships ShapeLearn GGUF quants for Qwen 3.8 27B

infoModel releaseimportance 62
What's new: A new MarkTechPost report (2026-09-18T08:40Z) fills in the official details of the Qwen3.8-Omni-Flash announcement that the previous summary (2026-09-18T07:09Z) noted lacked specifics: it is API-only with a 1M-token context window on the Qwen3.8-Flash-Next architecture, claims a 25%+ average gain over Qwen3.5-Omni-Plus across 29 evaluations (OmniVideoBench 63.4 to 67.8 at ~45.7% fewer tokens),…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

Alibaba's Qwen team announced Qwen3.8-Omni-Flash, an API-only omni-modal model with a 1M-token context window, agentic audio-video understanding, and tool use, claiming a 25%+ average improvement over Qwen3.5-Omni-Plus across 29 evaluations; no open weights…

Two developments hit the Qwen 3.8 family within roughly a day. Alibaba's Qwen team announced Qwen3.8-Omni-Flash (also written 'Qwen 3.8 Omni Flash') via its official blog, surfacing on Hacker News with 45 points and 7 comments. Follow-up details from MarkTechPost describe it as an API-only omni-modal model built on the Qwen3.8-Flash-Next architecture that accepts text, images, audio, and video and returns text, with a 1M-token context window and thinking enabled by default. It is positioned for fast, efficient multimodal inference: Qwen claims a 25%+ average improvement over Qwen3.5-Omni-Plus across 29 evaluations, with OmniVideoBench rising from 63.4 to 67.8 while using about 45.7% fewer tokens via coarse-to-fine agentic perception that gathers evidence instead of processing long video linearly. Qwen claims audio-visual performance near Gemini 3.8 Flash and audio performance above it. Audio input is supported in 113 languages, video up to 2 hours/2 GB, and availability spans 6 regions. The model is hosted on QwenCloud, Model Studio, and Qwen Studio at $0.15/$0.47 per 1M input/output tokens. No open weights were released, but Qwen open-sourced Qwen-MM-Plugins under Apache-2.0, adding omni-memory, omni-video2note, and omni-chatcut skills for agent harnesses. In parallel, ByteShape released its full ShapeLearn GGUF quantization set for Qwen 3.8 27B (base model released August 14, 2026), following earlier ShapeLearn-Lite quants published four days after launch. Five quants spanning IQ2_XXS 2.56bpw to IQ4_XS 3.84bpw were benchmarked on six GPUs against Unsloth Dynamic v3, ISTA-DASLab, Bartowski, and AtomicChat; the GPU-5 IQ4_XS quant reaches 99.63% of the BF16 aggregate score at roughly 90 tok/s on RTX Pro 6000 and RTX 5090, at 13.1 GB VRAM. Each GGUF bundles an MTP draft head, and a separate 1.1 GB DFlash2 draft model enables faster text-only speculative decoding via llama.cpp. The reports cover different models within the Qwen 3.8 family and do not conflict.

  • Alibaba's Qwen team announced Qwen3.8-Omni-Flash via its official blog on 2026-09-17; it drew 45 points and 7 comments on Hacker News
  • Qwen3.8-Omni-Flash is API-only, built on the Qwen3.8-Flash-Next architecture, accepts text/images/audio/video and returns text, with a 1M-token context window and thinking enabled by default
  • Qwen claims 25%+ average improvement over Qwen3.5-Omni-Plus across 29 evaluations; OmniVideoBench rose from 63.4 to 67.8 with ~45.7% fewer tokens via coarse-to-fine agentic perception
  • Qwen claims audio-visual performance near Gemini 3.8 Flash and audio performance above it
  • Audio input in 113 languages; video up to 2 hours/2 GB; available in 6 regions
  • Pricing: $0.15/$0.47 per 1M input/output tokens on QwenCloud, Model Studio, and Qwen Studio
  • No open weights for the model; Qwen open-sourced Qwen-MM-Plugins under Apache-2.0 with omni-memory, omni-video2note, and omni-chatcut skills
  • ByteShape released full ShapeLearn GGUF quants for Qwen 3.8 27B (base model released August 14, 2026), following ShapeLearn-Lite quants published four days after launch

Coverage timeline

  1. · 16h ago
    Hacker News · AI· 55
    Alibaba releases Qwen 3.8 Omni Flash

    Alibaba's Qwen team releases Qwen 3.8 Omni Flash, a fast omni-modal model announced on the official Qwen blog.

  2. · 13h ago
    Hacker News · AI· 30
    Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

    ByteShape released full ShapeLearn GGUF quants of Qwen 3.8 27B; its GPU-5 IQ4_XS reaches 99.63% of BF16 quality at 13.1 GB VRAM.

  3. · 6h ago
    MarkTechPost· 62
    Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use

    Alibaba's Qwen team launched Qwen3.8-Omni-Flash, an API-only omni-modal model with 1M-token context and agentic audio-video understanding and tool use.