orcarouter/OrcaSAQ-2-27B — new model trending #30 on Hugging Face
OrcaRouter released OrcaSAQ2 27B, a 12.3 GB 3-bit quant of Qwen3.8-27B with near-BF16 perplexity.
OrcaRouter released OrcaSAQ2 27B, a sensitivity-aware mixed-precision quantization of Qwen3.8-27B that shrinks a 54 GB BF16 checkpoint to 12.3 GB at about 3.21 bits per weight. On WikiText-2 it reports perplexity 5.6482 versus 5.6468 for BF16, a 0.02% increase, plus 93.2% top-1 agreement and 0.031 mean KLD. The Apache-2.0 checkpoint keeps a 262K context window, thinking mode, and tool calling, and is served with vLLM. The card lists 70.0 on SWE-bench Verified and 58.4 on Terminal-Bench 2.1, noting those figures are not strict model-only comparisons.
- 54 GB BF16 checkpoint compressed to 12.3 GB, about 77% smaller.
- Perplexity rises only 0.02%, with 93.2% top-1 agreement.
- Supports 262K context, thinking mode, tool calling, and vLLM.
- Card cites 70.0 SWE-bench Verified and 58.4 Terminal-Bench 2.1.
Full article1,766 words · extracted from huggingface.co · click to collapse
<p align="center">
<a href="https://www.orcarouter.ai">
<img src="https://www.orcarouter.ai/orca-logo-classic.png" width="96" alt="OrcaRouter">
</a>
</p>
<h1 align="center">OrcaSAQ2 27B</h1>
<p align="center"><strong>High-fidelity 3-bit Qwen3.8 for long-horizon agents.</strong></p>
<p align="center"><strong>54 GB → 12.3 GB · +0.02% PPL · 93.2% Top-1 Agreement · 0.031 KLD · 262K Context</strong></p>
<p align="center">
<a href="https://www.orcarouter.ai"><strong>OrcaRouter AI Gateway</strong></a> ·
<a href="https://x.com/OrcaRouter"><strong>X</strong></a> ·
<a href="https://discord.gg/yAh6Tex6kx"><strong>Discord</strong></a> ·
<a href="https://github.com/Continuum-AI-Corp"><strong>GitHub</strong></a> ·
<a href="https://www.orcarouter.ai/models"><strong>All Models</strong></a>
</p>
---
> ## 27B reasoning. 12.3 GB.
>
> **OrcaSAQ2 27B** compresses Qwen3.8-27B from a **54 GB BF16 checkpoint to 12.3 GB** while preserving extremely high fidelity to the original model.
>
> Built for: **long-horizon agents · coding · tool use · reasoning · stateful execution**
OrcaSAQ2 is a proprietary **sensitivity-aware mixed-precision quantization system** developed by OrcaRouter and its research team behind.
It is optimized around one goal: **Preserve as much useful model behavior as possible inside a practical GPU memory envelope.**
The resulting checkpoint provides:
- **77.2% smaller storage footprint**
- only **+0.02% perplexity** versus BF16
- **93.2% token-level Top-1 agreement**
- **0.031 mean KLD**
- **262K context**
- thinking mode
- tool calling
- MTP speculative decoding
- production serving through **vLLM**
---
## At a glance
| Metric | BF16 | OrcaSAQ2 |
|---|---:|---:|
| Checkpoint | 54 GB | **12.3 GB** |
| Relative size | 100% | **22.8%** |
| Storage reduction | — | **77.2%** |
| Decoder precision | 16-bit | **3.21 bpw avg.** |
| Perplexity | 5.6468 | **5.6482** |
| PPL delta | — | **+0.02%** |
| Top-1 agreement | 100% | **93.2%** |
| Mean KLD | — | **0.031** |
| Context | 262K | **262K** |
<p align="center"><strong>4.4× smaller. +0.02% perplexity.</strong></p>