ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF — new model trending #30 on Hugging Face
ISTA-DASLab published a 50-percent expert-pruned Qwen3.8 coder weighing 58.4 GB in GGUF.
ISTA-DASLab released Qwen3.8-Flash-Next-GSQ-RCO-Coder, a GGUF compression of the 176.9-billion-parameter Qwen3.8-Flash-Next mixture-of-experts model that removes half of its 512 routed experts and stores remaining weights at 3.5 bits. The package is 58.4 GB versus 354 GB at BF16, an effective 1.89 bits per original parameter, with 29.6 GB required in memory. At xhigh reasoning effort, SWE-bench Verified is 75.60 versus 82.80 for the base, and LiveCodeBench v6 is 86.28 versus 87.43. Expert selection uses RCO to minimize KL divergence under a uniform per-layer expert count required by GGUF. The release is Apache 2.0 and was trending at number 30 on Hugging Face.
- Half of 512 routed experts are removed; retained weights stay at 3.5 bits.
- The GGUF package is 58.4 GB versus a 354 GB BF16 base.
- SWE-bench Verified is 75.60 versus 82.80, retaining 91.3% of the base.
- LiveCodeBench v6 is 86.28 versus 87.43, retaining 98.7%.
- About 29.6 GB must stay in memory, targeting one 32 GB accelerator.
Full article1,408 words · extracted from huggingface.co · click to collapse
<div align="center">
<a href="https://github.com/IST-DASLab"><img src="assets/banner.png" alt="GGUF, GSQ-RCO expert-pruned coder" width="100%"/></a>
<br/>
# Qwen3.8-Flash-Next · GSQ-RCO Coder
**A 512-expert mixture-of-experts model with half of its experts removed**, targeted at code and retaining multimodal capability.
**58.4 GB** against 354 GB at BF16, an effective **1.89 bits per parameter of the original transformer**.
[](https://arxiv.org/abs/2604.18556)
[](https://arxiv.org/abs/2605.00649)
[](https://github.com/IST-DASLab/GSQ)
[](https://github.com/IST-DASLab/RCO)
[](https://github.com/IST-DASLab)
[](#license)
</div>

---
## Overview
This release is a capability-targeted compression of Qwen3.8-Flash-Next. Half of the routed experts are removed from the model rather than quantized, and the remaining weights are stored at 3.5 bpw.
The base model has 176.9B parameters and occupies 354 GB at BF16. The compressed model reduces to only 29.6 GB which must reside in memory: the n-gram shard is a lookup table and may be served from disk. The resident working set of a 176.9B-parameter model is therefore within the capacity of a single 32 GB accelerator.
Averaged over the transformer, this corresponds to **1.89 bits per parameter of the original model**. The figure is an effective rate: it amortises the removed experts over the original parameter count, and so expresses the combined effect of pruning and quantization. No individual weight is stored at 1.89 bits. The retained weights are stored at 3.5 bpw, unchanged by pruning, and the halved expert count accounts for the remainder of the reduction. Following the convention of the accompanying releases, the figure is reported over the transformer and excludes the fixed-precision n-gram shard.
Removing half of a model's experts necessarily reduces its capabilities. The contribution of this release is that the reduction is directed towards code, agentic tool use, vision, and spatial reasoning, and the retained experts are therefore those that support these domains. Degradation outside this set is an accepted cost of the method. For general-purpose use, the [unpruned GSQ-RCO releases](https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF) are the appropriate choice.