Alibaba releases 7B Qwen-Image-2.1 as community quantizations spread
Alibaba released open-weight 7B Qwen-Image-2.1 for image generation and editing; community GGUF, uncensored, and distilled LoRA builds are now trending.
Alibaba’s Qwen team released Qwen-Image-2.1, an open-weight text-to-image and editing model whose diffusion transformer has 7 billion parameters in 32 DiT layers, down from the 20B Qwen-Image model shipped in August 2025. It loads an 8B Qwen3-VL encoder and a 64-channel RGBA VAE, supports native transparency, up to 10 reference images, and local edits, defaults to 2048×2048, and is offered on Hugging Face, GitHub, and Model Scope under a research license that bars commercial use without a separate agreement; The Decoder adds that it runs on consumer GPUs such as the RTX 3090, with day-0 support in Diffusers, ComfyUI, vLLM-Omni, and SGLang. Benchmark claims diverge: The Decoder says Alibaba claims the model beats most closed models on its own benchmark, with independent evaluations still pending, while MarkTechPost reports 60.28 on Qwen-Image-Bench, above FLUX 2 Max at 55.33 but below GPT Image 2.5 Sunburst at 67.01. Community Hugging Face uploads followed, including abenzerps GGUF quantizations marketed as uncensored and without safety filters—one repo cited Q4_K_M through Q8_0, another Q4_0 at 4.05 GB to Q8_0 at 7.59 GB plus a Q8_0 tensor-shape bug in ComfyUI—plus pottokao Heretic text-encoder builds and Unsloth Dynamic 2.0 GGUFs of the denoiser only. Sources slightly differ on the text encoder’s BF16 size (17.53 GB versus 17.5 GB); Unsloth cites a Q4_K_XL result of LPIPS 0.029, SSIM 0.959, 5.15 GB, and 36.5 seconds versus Q4_K_M. On 22 September 2026 Viggle released viggle-turbo v0.2.1, a six-step rank-256 LoRA (1.3 GB; a rank-128 cut is 680 MB) with no classifier-free guidance that is about five times faster than the 40-step base on one to three reference images, reporting 0.98 times base diversity and 0% composition drift on 96 held-out requests, while dense text and complex multi-reference edits can still lag.
- Alibaba’s Qwen team released open-weight Qwen-Image-2.1: a 7B-parameter diffusion transformer with 32 DiT layers, down from 20B Qwen-Image (August 2025), plus an 8B Qwen3-VL encoder and a 64-channel RGBA VAE.
- One checkpoint does generation, editing, native RGBA, and up to 10 reference images, defaulting to 2048×2048; a research license bars commercial use without a separate agreement. Hosted on Hugging Face, GitHub, and Model Scope, with…
- Benchmark wording disagrees: The Decoder says Alibaba claims it beats most closed models on its own bench, with independent evals pending; MarkTechPost reports Qwen-Image-Bench 60.28, above FLUX 2 Max at 55.33 but below GPT Image 2.5…
- On 20 Sep 2026, abenzerps GGUF repos trended, including an explicitly uncensored build with no safety filters. Cited ranges differ: Q4_K_M to Q8_0 versus Q4_0 at 4.05 GB to Q8_0 at 7.59 GB, with a reported Q8_0 ComfyUI tensor-shape bug.
- Text-encoder BF16 size is 17.53 GB in one report and 17.5 GB in pottokao’s Heretic Qwen3-VL-8B builds, which also list Q4_K_M at 5.0 GB plus a 1.2 GB mmproj and FP8 at 9.3 GB.
- Unsloth’s Dynamic 2.0 GGUF covers only the denoiser and cites Q4_K_XL at LPIPS 0.029, SSIM 0.959, 5.15 GB, and 36.5 s versus Q4_K_M; it still needs a separate BF16 VAE and Qwen3-VL-8B encoder.
- Viggle viggle-turbo v0.2.1 (22 Sep 2026) is a 6-step LoRA with no CFG, about 5× faster than the 40-step base: rank-256 is 1.3 GB and rank-128 is 680 MB. On 96 held-out requests it shows 0.98× diversity and 0% composition drift, but dense…
Coverage timelineoldest first · each row is one article
- · 6d agoabenzerps/Qwen-Image-2.1-Uncensored-GGUF — new model trending #30 on Hugging Face
Hugging Face trending models· 35
An uncensored GGUF version of Qwen-Image-2.1 for local image generation is trending, featuring no content filters and optimized for ComfyUI.
- · 6d agoabenzerps/Qwen-Image-2.1-GGUF — new model trending #14 on Hugging Face
Hugging Face trending models· 45
GGUF quantizations of Qwen-Image-2.1 image generation model released for local inference with multiple quantization options.
- · 6d ago