ZeroHour
The Decoderpublished ()ingested Matthias Bastian
Part of a story covered by 4 sources: “Alibaba's Qwen ships Qwen3.8-Omni-Flash, a 1M-context agent model undercutting Gemini Flash pricing, as ByteShape quantizes Qwen 3.8 27B” — merged summary and timeline →

Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks

infoModel releaseimportance 65
AI summary · glm-5.3-flash

Qwen launches Qwen3.8-Omni-Flash, a multimodal agent model priced far below Google's Gemini 3.8 Flash with similar benchmark performance.

Qwen3.8-Omni-Flash is Qwen's first multimodal model built for AI agents, jointly processing audio and video with tool use for tasks like vlog editing and movie summarization. It offers a one-million-token context window, and Qwen claims it roughly matches Gemini 3.8 Flash on multimodal benchmarks. API pricing is $0.15 per million input tokens and $0.47 per million output tokens, versus Gemini 3.8 Flash's introductory $0.75 input and $3.75 output rates. The model is available via Qwen Studio, Qwen Cloud, and the API, with open-source Qwen-MM-Plugins and a Qwen-Live Harness for real-time camera and microphone interaction.

  • First Qwen multimodal model designed specifically for autonomous AI agents
  • One-million-token context window with joint audio-video processing
  • Undercut pricing: $0.15/$0.47 per million tokens vs Gemini Flash $0.75/$3.75
  • Qwen-MM-Plugins add video editing and speaker recognition to Claude Code and Gemini CLI
Full article235 words · extracted from the-decoder.com · click to collapse

Qwen3.8-Omni-Flash is Qwen's first multimodal model built for AI agents. It processes audio and video together, draws conclusions, and uses tools on its own to edit vlogs, translate short videos, or summarize movies. The context window spans one million tokens. On audio-video tasks, Qwen says it comes close to matching Gemini 3.8 Flash.

Qwen 3.8 Omni Flash performs on par with Gemini Flash 3.8 in multimodal benchmarks but is much more affordable. | Image: Qwen

API pricing sits at $0.15 per million input tokens and $0.47 per million output tokens. Qwen estimates audio input at under $0.01 per hour, while 720p video with audio at one frame per second runs about $0.20, not counting response costs. For comparison, Gemini 3.8 Flash charges $0.75 for input and $3.75 for output per million tokens at its introductory rate, with prices set to double on January 1, 2027.

The model is available through Qwen Studio, Qwen Cloud, and the API. The open-source Qwen-MM-Plugins add video editing, speaker recognition, PDF video notes, and reusable workflows to agents like Claude Code, Gemini CLI, and Qwen Code. Qwen-Live Harness enables real-time interaction using a camera and microphone.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now

Text extracted automatically; images, tables and formatting may be missing. Original: https://the-decoder.com/qwen3-8-omni-flash-undercuts-gemini-flash-pricing-while-matching-its-multimodal-benchmarks/