autotrust published GEV-26B-Decide-NVFP4, an NVFP4 expert-quantized 26B model that fits in about 17 GB.
autotrust released GEV-26B-Decide-NVFP4, quantizing the 3,840 routed expert MLPs of GEV-26B-Decide to NVIDIA NVFP4 while leaving attention, routers, the vision tower, and decision heads in bf16. Download size falls from 49.5 GB to 18 GB, and vLLM weight memory from 51.1 GiB to 17.1 GiB. On System 1, GPQA Diamond moves from 43.9% to 44.9% and HLE text multiple choice from 8.4% to 9.7%, within run-to-run noise. The bf16 model scores 62.48 on Decision Index 0.2.1 with adaptive thinking.
3,840 routed expert MLPs quantized to NVFP4; other weights stay bf16.
vLLM weight memory falls from 51.1 GiB to 17.1 GiB.
System 1 scores 44.9% on GPQA Diamond and 9.7% on HLE text.
Native W4A4 on Blackwell; Marlin W4A16 on Hopper and Ampere.
bf16 GEV-26B-Decide scores 62.48 on Decision Index 0.2.1.
Full article3,281 words · extracted from huggingface.co · click to collapse
# autotrust/GEV-26B-Decide-NVFP4
### [autotrust/GEV-26B-Decide](https://huggingface.co/autotrust/GEV-26B-Decide) with NVFP4 routed experts: the same model in 17 GB of GPU memory instead of 50 GB
## NVFP4 version
This repository is **autotrust/GEV-26B-Decide with its 3,840 routed expert MLPs (30 layers × 128 experts) quantized
to NVFP4**. Everything else is the bf16 model unchanged: attention, the dense MLP of every layer, the routers, the
vision tower, the lm_head, the System 1 adapter, the decision head and the temperatures. The rest of this card is the
card of GEV-26B-Decide; its benchmark tables are the bf16 results, and the NVFP4 comparison is below.
**GEV-26B-Decide answers typed questions with a calibrated probability for every option; with thinking switched on, it
thinks only when it needs to.** System 1 decides in one forward pass (about 45 ms). When its leading option is uncertain, System 2 (the same
backbone in Gemma-4 thinking mode) reasons over the question, and the reasoning is folded into the final probabilities.
One set of weights, one vLLM engine, for text and images.
| | what it does | output |
|---|---|---|
| **System 1** | typed decisions: yes/no · pick one of 2–256 options · rate 0–5, over text and images; prompts up to 256K tokens | a calibrated probability for every option, in one forward pass |
| **Adaptive thinking** (opt-in: `thinking: "auto"`) | System 1 first; below 0.8 confidence, System 2 thinks and its answer is folded in | calibrated probabilities |
| **System 2** | the unmodified `google/gemma-4-26B-A4B-it`, optionally thinking step by step, text and images | text / reasoning |
GEV-26B-Decide was previously published as `autotrust/JEV-Gemma4-26B-A4B`; the weights are the same.
> **Two models, two organisations.** **TypeSafe Jev 1.13** is the hosted, closed model made by TypeSafe AI.
> **autotrust/GEV-26B-Decide** is an independent open-weights model built by AutoTrust AI; it is not affiliated with,
> endorsed by, or a product of TypeSafe AI.
## Watch it think
Each puzzle has one correct answer. System 1 answers in one pass; with `thinking: "auto"`, System 2 thinks when System 1
is below 0.8 confidence, and its answer is folded into the probabilities. The videos show puzzles that System 1 got wrong;
the tables below count all puzzles.
**Minesweeper.** Which hidden cell is certainly safe? System 1 is at chance (21.0 % against 25 %); with thinking, 86.0 %.
These thoughts usually reach the 8,192-token budget, and the answer read at that point is still right most of the time.