Qwen-Image-2.1 open release followed by GGUFs and a faster LoRA
Alibaba’s 7B Qwen-Image-2.1 unifies generation and editing; community GGUFs and a six-step LoRA followed, while a later audio release is separate.
Alibaba’s Qwen team announced Qwen-Image-2.1 on 20 September 2026, an open-weight model whose diffusion transformer has 7 billion parameters across 32 single-stream DiT layers, down from the 20B Qwen-Image shipped in August 2025. One checkpoint also loads an 8B Qwen3-VL encoder and a 64-channel RGBA VAE, defaults to 2048×2048, natively handles transparency, accepts up to 10 reference images, and is described as running on consumer GPUs such as the RTX 3090; weights are on Hugging Face, GitHub, and ModelScope under a research license that bars commercial use without a separate agreement, with day-0 support in Diffusers, ComfyUI, vLLM-Omni, and SGLang. Sources disagree on quality: The Decoder says Alibaba claims the model beats most closed models on its own benchmark and that independent evaluations are pending, while MarkTechPost reports 60.28 on Qwen-Image-Bench, above FLUX 2 Max at 55.33 but below GPT Image 2.5 Sunburst at 67.01. Community weights followed within days, including abenzerps GGUFs from Q4_0 (4.05 GB) to Q8_0 (7.59 GB) marketed without safety filters and with a known Q8_0 ComfyUI tensor-shape bug, Heretic text-encoder builds, and Unsloth Dynamic 2.0 denoiser GGUFs whose Q4_K_XL is listed at LPIPS 0.029, SSIM 0.959, 5.15 GB, and 36.5 seconds per render. Viggle’s v0.2.1 six-step LoRA, about five times faster than the 40-step base and using no classifier-free guidance, measured 0.98 times base diversity and 0% composition drift on 96 held-out requests, but it supports only one to three reference images and can lag on dense text and complex edits. A 23 September report on five Qwen-Audio-3.1 speech models and Qwen Cloud audio price cuts of up to 95% is a separate launch and does not change these image-model facts.
- Announced 20 September 2026: Qwen-Image-2.1 uses a 7B diffusion transformer with 32 single-stream DiT layers, down from the 20B Qwen-Image of August 2025, plus an 8B Qwen3-VL encoder and a 64-channel RGBA VAE, defaulting to 2048×2048.
- One checkpoint covers generation, editing, native transparency, and up to 10 reference images; it is described as running on an RTX 3090, with weights on Hugging Face, GitHub, and ModelScope under a research license that bars commercial…
- Day-0 tooling named in reports: Diffusers, ComfyUI, vLLM-Omni, and SGLang.
- Benchmark claims differ: The Decoder says Alibaba claims it beats most closed models on its own benchmark, with independent evaluations pending; MarkTechPost reports 60.28 on Qwen-Image-Bench, above FLUX 2 Max at 55.33 and below GPT Image…
- Community GGUFs include abenzerps builds from Q4_0 (4.05 GB) to Q8_0 (7.59 GB), marketed without safety filters, with a known Q8_0 ComfyUI tensor-shape bug, and Unsloth Dynamic 2.0 Q4_K_XL at LPIPS 0.029, SSIM 0.959, 5.15 GB, and 36.5…
- Viggle v0.2.1 is a six-step LoRA, about 5× faster than the 40-step base with no classifier-free guidance; the rank-256 adapter is 1.3 GB and rank-128 is 680 MB, with 0.98× base diversity and 0% composition drift on 96 held-out requests,…
- A separate 23 September report covers Qwen-Audio-3.1, five ASR, TTS, and realtime speech models, with Qwen Cloud cuts of about 70% for TTS, about 85% for Realtime, and up to 95% for ASR.
Coverage timelineoldest first · each row is one article
- · 6d agoQwen-Image-2.1: Compact, efficient, and unified image creation
Hacker News · security· 45
Alibaba's Qwen team released Qwen-Image-2.1, a compact open-source unified text-to-image generation and image editing model.
- · 6d agoabenzerps/Qwen-Image-2.1-Uncensored-GGUF — new model trending #30 on Hugging Face
Hugging Face trending models· 35
An uncensored GGUF version of Qwen-Image-2.1 for local image generation is trending, featuring no content filters and optimized for ComfyUI.
- · 6d agoabenzerps/Qwen-Image-2.1-GGUF — new model trending #14 on Hugging Face
Hugging Face trending models· 45