ConwayResearch/Underdog-Saluki-27B-1.0 — new model trending #30 on Hugging Face
Underdog Saluki 27B 1.0 is a 7.89 GB GGUF of Qwen3.8-27B that retains tool calling for llama.cpp.
Conway Research published Underdog-Saluki-27B-1.0, a 7.89 GB IQ2-mix GGUF derived from Qwen3.8-27B for stock llama.cpp. On a 120-task Underdog Bench subset of BFCL v4, it scored 88 versus 84 for the full 54 GB model, and 42 versus 35 on 100 parallel tool-call tasks, with about 96% average retention across nine benchmarks. Competition-math retention is roughly 82 to 85 percent; a public AIME 2025 figure cited for the full model is 96.7 versus 79.2 here, from a different harness. An optional mmproj file adds vision, and the weights are Apache 2.0.
- Saluki is a 7.89 GB IQ2-mix GGUF of Qwen3.8-27B for stock llama.cpp.
- Tool-calling score is 88 of 120 versus 84 for the full 54 GB model.
- Parallel calls scored 42 versus 35, with about 96 percent average benchmark retention.
- Math and some coding scores drop; vision needs a separate 629 to 928 MB file.
- Released under Apache 2.0 by Underdog, based on Qwen and ISTA-DASLab GGUF work.
Full article570 words · extracted from huggingface.co · click to collapse
# Underdog Saluki 27B 1.0
Qwen3.8-27B in under 8 GB, tuned to keep tool calling intact. A standard GGUF for stock llama.cpp.
| Tool calling (Underdog Bench) | Parallel tool calls | Average retention | Size |
|---|---|---|---|
| **88** vs 84 for the full model | **42** vs 35 for the full model | **96%** across 9 benchmarks | **7.89 GB** vs 54 GB |
| | |
|---|---|
| Base model | Qwen3.8-27B |
| File | `Underdog-Saluki-27B-1.0-IQ2-mix.gguf`, 7.89 GB |
| Runtime | stock llama.cpp, and apps built on it |
| Modality | text, plus images with the optional vision add-on |
| Vision add-on | `mmproj-Underdog-Saluki-27B-1.0-F16.gguf` (928 MB) or `-Q8_0.gguf` (629 MB) |
| Thinking | on by default, switchable per request, same as Qwen3.8 |
## Quickstart
```bash
huggingface-cli download ConwayResearch/Underdog-Saluki-27B-1.0 Underdog-Saluki-27B-1.0-IQ2-mix.gguf --local-dir .
```
```bash
llama-server -m Underdog-Saluki-27B-1.0-IQ2-mix.gguf --jinja -ngl 99 -fa on -c 32768
```
**Add vision.** Download the vision add-on too, and pass it with `--mmproj`. Then send images through the OpenAI chat API as usual.
```bash
huggingface-cli download ConwayResearch/Underdog-Saluki-27B-1.0 mmproj-Underdog-Saluki-27B-1.0-F16.gguf --local-dir .
```
```bash
llama-server -m Underdog-Saluki-27B-1.0-IQ2-mix.gguf --mmproj mmproj-Underdog-Saluki-27B-1.0-F16.gguf --jinja -ngl 99 -fa on -c 32768
```
`--jinja` turns on the Qwen3.8 chat template, which handles tool calls and thinking. The server speaks the OpenAI chat API on port 8080.
| Use | Settings |
|---|---|
| General use, reasoning, instructions | thinking on (default), temperature 0.6, top_p 0.95, top_k 20 |
| Fast, direct tool calls | thinking off with `"chat_template_kwargs": {"enable_thinking": false}`, temperature 0 |
## Benchmarks
Saluki ran in stock llama.cpp. Where we ran the full-size Qwen3.8-27B ourselves, it used the same harness and settings.
**Tool calling.** Underdog Bench is 120 tasks from the Berkeley Function Calling Leaderboard (BFCL v4), frozen before any model was tested. Thinking off, temperature 0.
| Model | Size | Passed (of 120) |
|---|---|---|
| **Underdog Saluki 27B 1.0** | 7.89 GB | **88** |
| Qwen3.8-27B, full size | 54 GB | 84 |
| Bonsai 2 | 5.95 GB | 70 |