ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF — new model trending #30 on Hugging Face
UkisAI published GSQ-RCO GGUF quantizations of Swift 1.5 Qwen3.8-27B, from 8.42 GB to 11.77 GB.
UkisAI published mixed-precision GGUF quantizations of Swift 1.5 Qwen3.8-27B, using per-tensor allocations from ISTA-DASLab's GSQ-RCO release. Standard files are IQ2_XS (8.42 GB, development KLD 0.189979), IQ2_S (9.26 GB, 0.134751), IQ3_XXS (10.09 GB, 0.097774), and IQ3_S (11.77 GB, 0.051265); optional MTP-head variants are about 0.35 GB larger. The card says Swift 1.5 uses 58.5% fewer thinking tokens and scores 0.35% higher than its base, for a 9.18× speed-up on several tasks; those figures are separate from the quantization KLD tests, which use a 512-token context and are not task-accuracy scores. All four refined files beat their matched Swift starting quants on seven held-out domains, but they are not uniformly better than the ISTA comparison quants.
- IQ2_XS through IQ3_S GGUFs range from 8.42 GB to 11.77 GB.
- Optional MTP files are about 0.35 GB larger and need a compatible runtime.
- KLD is versus Swift 1.5 BF16 on 512-token chunks, not task scores.
- Card claims 58.5% fewer thinking tokens and a 9.18× speed-up versus the base.
- llama.cpp with Qwen3.8 support is required; no vision projector is included.
Full article947 words · extracted from huggingface.co · click to collapse
<div align="center">
<a href="https://ukisai.com"><img src="ukisai-banner.png" alt="UkisAI" style="width:100%;height:auto;" /></a>
<p><a href="https://ukisai.com"><b>Website</b></a> • <a href="https://ukisai.com/products/swift"><b>Learn more</b></a> • <a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b"><b>BF16 model</b></a> • <a href="https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GGUF"><b>Standard GGUFs</b></a> • <a href="#evaluation"><b>Evaluation</b></a> • <a href="#license-and-access"><b>Enterprise licensing</b></a></p>
</div>
# Swift 1.5 Qwen3.8-27B · GSQ-RCO
Compact, mixed-precision GGUF quantizations of [Swift 1.5 Qwen3.8-27B](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b), with Swift-specific refinement using the per-tensor allocations from [ISTA-DASLab's GSQ-RCO release](https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF).
Swift 1.5 builds on Swift 1.0 through post-training focused on long-horizon, agentic and coding tasks, improving overall performance while using fewer thinking tokens. See the [original model card](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b) for the model's training approach and benchmark results. Those model-level benchmarks are separate from the quantization measurements below.
Swift 1.5 uses **58.5% fewer thinking tokens** while scoring **0.35% higher** than the base, for a **9.18× speed-up** on several tasks.
## Available quantizations
Each tier is a single GGUF file. Sizes are decimal GB; runtime memory also includes the context cache and compute buffers. Tier names denote mixed-precision allocation profiles, rather than a uniform type for every tensor.
| Tier | Standard GGUF | With MTP head | Development KLD ↓ |
| --- | ---: | ---: | ---: |
| IQ2_XS | [8.42 GB](Swift-1.5-Qwen3.8-27B-GSQ-RCO-IQ2_XS.gguf) | [8.77 GB](Swift-1.5-Qwen3.8-27B-GSQ-RCO-IQ2_XS-mtp.gguf) | 0.189979 |
| IQ2_S | [9.26 GB](Swift-1.5-Qwen3.8-27B-GSQ-RCO-IQ2_S.gguf) | [9.61 GB](Swift-1.5-Qwen3.8-27B-GSQ-RCO-IQ2_S-mtp.gguf) | 0.134751 |
| IQ3_XXS | [10.09 GB](Swift-1.5-Qwen3.8-27B-GSQ-RCO-IQ3_XXS.gguf) | [10.44 GB](Swift-1.5-Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf) | 0.097774 |
| IQ3_S | [11.77 GB](Swift-1.5-Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf) | [12.12 GB](Swift-1.5-Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf) | 0.051265 |
The optional `-mtp` files retain the matching refined model tensors and add the MTP head. They require a runtime with support for this model's MTP implementation. The KLD results here were measured on the **standard files**; MTP decoding speed and quality have not been separately evaluated.