Skip to content
autotrust/GEV-26B-Decide — new model trending #30 on Hugging Face · ZeroHour 01:58 UTC · 18m ago
today 36K visits · 35K views · 6 reads · all time 1.2M views Hugging Face trending models· published Oct 2, 04:20 UTC (6d ago ) · ingested Oct 5, 14:17 UTC · by autotrust Part of a story covered by 2 sources: “AutoTrust releases GEV-26B-Decide and NVFP4 variant” — merged summary and timeline →autotrust/GEV-26B-Decide — new model trending #30 on Hugging Face AI summary · grok-4.7
AutoTrust released GEV-26B-Decide, an open 26B Gemma-4 decision model scoring 62.48 on Decision Index 0.2.1.
AutoTrust AI published autotrust/GEV-26B-Decide, open weights previously named JEV-Gemma4-26B-A4B, built on Google's Gemma-4-26B-A4B-it (26 billion parameters, about 4 billion active). On Decision Index 0.2.1 it scores 62.48 balanced skill, above the TypeSafe Jev 1.13 board mark of 57.91; area scores include Tools and Automation 0.697 and Arts and Taste 0.415. System 1 returns calibrated option probabilities in one pass, about 45 ms on a B200, and adaptive thinking invokes Gemma-4 reasoning below 0.8 confidence. In demos it finished 95% of 60 browser tasks at about 85 ms per click, but only 40% of 20 robot-arm pick-and-place scenes.
26B MoE on Gemma-4-26B-A4B-it, about 4B active parameters. Decision Index 0.2.1 score is 62.48 versus TypeSafe Jev 1.13 at 57.91. Computer-use demo: 95% of 60 tasks, about 85 ms per click. Robot-arm pick-and-place: 40% of 20 scenes at 61 ms per step. Same weights as earlier JEV-Gemma4-26B-A4B; not a TypeSafe product. Full article 3,254 words · extracted from huggingface.co · click to collapse # autotrust/GEV-26B-Decide
### A decision model with adaptive thinking, on Gemma-4-26B-A4B-it (26 B parameters, ≈ 4 B active per token) — Decision Index 0.2.1: **62.48**
## Decision Index
| | Decision Index 0.2.1 (balanced skill) | balanced raw | breadth skill |
|---|---:|---:|---:|
| **autotrust/GEV-26B-Decide, adaptive thinking** | **62.48** | 70.66 | 62.00 |
| TypeSafe Jev 1.13 (board) | 57.91 | — | — |
| area (skill) | Knowledge & Reasoning | Language | Retrieval & Classification | Tools & Automation | Arts & Taste |
|---|---:|---:|---:|---:|---:|
| GEV-26B-Decide, adaptive thinking | **0.602** | 0.636 | 0.679 | 0.697 | 0.415 |
How the score was computed (our scoring with the kit's `score --edition 0.2.1`, not a board entry):
* **Knowledge & Reasoning:** all ten benchmarks with adaptive thinking (per-benchmark table under [Adaptive thinking](#adaptive-thinking)).
* **Other four areas:** System 1 only (thinking off), from the complete System 1 run of these weights (all 150,759
requests, 0 errors), results in
[`autotrust/jev-decision-index-results`](https://huggingface.co/datasets/autotrust/jev-decision-index-results)
(`runs/jev-gemma4-26b-a4b`, the weights' previous name). Thinking was tried on five of their benchmarks and is not used
there: it adds little to classification, retrieval and tool selection (see
[Outside Knowledge & Reasoning](#outside-knowledge--reasoning)).
* The board requires a median latency of at most 1,000 ms per request, measured on an RTX PRO 6000. System 1 answers in
about 45 ms on a B200; adaptive thinking is far slower on the Knowledge & Reasoning benchmarks (see
[Latency](#latency)).
Details: `reports/decision_index_adaptive.json`, `reports/adaptive_latency_summary.json`.
## New (3 October 2026): computer use and robot arm, 60–85 ms per step
The fastest vision model of the family. Every step below is **one System 1 decision** (thinking off): a screenshot or a
camera image in, a probability for every action out, in a single forward pass on one B200.
**Computer use: screenshot → which element to click.** A real browser (headless Chromium). Every clickable element gets a
numbered box; System 1 picks the next click (or "the task is complete"), the browser clicks it, and the loop repeats.
<video src="https://huggingface.co/autotrust/GEV-26B-Decide/resolve/main/videos/computer_use_shop.mp4" controls autoplay loop muted playsinline width="100%"></video>
Related stories shared CVEs / entities Category Model release
Classifier grok-4.7
Confidence 93%
Ingested 3d ago
**95% of 60 random multi-step tasks completed** (shop, settings, mail; 3–7 clicks each) in **about 85 ms per click**:
the same success rate as JEV-27B-VL, 3× faster. The colour swatches carry no text, so that click is decided from the
**Robot arm: pick and place from a camera image.** At every step System 1 looks at the top camera image and answers two
questions: is the target left or right of the gripper, and above or below it? The arm moves accordingly and halves its
step whenever an answer flips. It grasps the cube, carries it and drops it in the tray (MuJoCo simulation).
<video src="https://huggingface.co/autotrust/GEV-26B-Decide/resolve/main/videos/robot_arm_pick_place.mp4" controls autoplay loop muted playsinline width="100%"></video>
61 ms per decision, so a whole pick and place takes 4–8 seconds of model time. It completed 40% of 20 random scenes:
close to the target its left/right answers are less precise than JEV-27B-VL's, so more grasps miss. Once grasped, 8 of 9
Same scenes and tasks for every model in the family:
| | **GEV-26B-Decide** | [JEV-27B-VL](https://huggingface.co/autotrust/JEV-27B-VL) | [JEV-9B](https://huggingface.co/autotrust/JEV-9B) |
| computer use: numbered boxes + element text (60 tasks) | 95% | 95% | 95% |
| time per click | **≈ 85 ms** | ≈ 260 ms | ≈ 200 ms |
| robot arm: pick and place (20 scenes) | 40% | **75%** | 50% |
| time per robot-arm decision | **61 ms** | 239 ms | 163 ms |
* For computer use, give it the element text (as an accessibility tree would). With the numbered boxes alone it
completes 15% and often declares the task complete too early.
* Use System 1 for simple visual questions inside a control loop. Asked to pick one of 8 motor commands directly, it
completed 0 of 10 scenes.
Demo code: [JEV-9B `vl/demos/`](https://huggingface.co/autotrust/JEV-9B/tree/main/vl/demos) (set `JEV_URL` to this
server). Per-episode results: [`reports/demos/`](reports/demos).
**GEV-26B-Decide answers typed questions with a calibrated probability for every option; with thinking switched on, it
thinks only when it needs to.** System 1 decides in one forward pass (about 45 ms). When its leading option is uncertain, System 2 (the same
backbone in Gemma-4 thinking mode) reasons over the question, and the reasoning is folded into the final probabilities.
One set of weights, one vLLM engine, for text and images.
| | what it does | output |
| **System 1** | typed decisions: yes/no · pick one of 2–256 options · rate 0–5, over text and images; prompts up to 256K tokens | a calibrated probability for every option, in one forward pass |
| **Adaptive thinking** (opt-in: `thinking: "auto"`) | System 1 first; below 0.8 confidence, System 2 thinks and its answer is folded in | calibrated probabilities |
| **System 2** | the unmodified `google/gemma-4-26B-A4B-it`, optionally thinking step by step, text and images | text / reasoning |
GEV-26B-Decide was previously published as `autotrust/JEV-Gemma4-26B-A4B`; the weights are the same.
> **Two models, two organisations.** **TypeSafe Jev 1.13** is the hosted, closed model made by TypeSafe AI.
> **autotrust/GEV-26B-Decide** is an independent open-weights model built by AutoTrust AI; it is not affiliated with,
> endorsed by, or a product of TypeSafe AI.
Each puzzle has one correct answer. System 1 answers in one pass; with `thinking: "auto"`, System 2 thinks when System 1
is below 0.8 confidence, and its answer is folded into the probabilities. The videos show puzzles that System 1 got wrong;
the tables below count all puzzles.
**Minesweeper.** Which hidden cell is certainly safe? System 1 is at chance (21.0 % against 25 %); with thinking, 86.0 %.
These thoughts usually reach the 8,192-token budget, and the answer read at that point is still right most of the time.
<video src="https://huggingface.co/autotrust/GEV-26B-Decide/resolve/main/videos/think_minesweeper.mp4" controls autoplay loop muted playsinline width="100%"></video>
**Connect Four.** Which column wins now, or stops the opponent from winning next move? 54.0 % → 99.5 %.
<video src="https://huggingface.co/autotrust/GEV-26B-Decide/resolve/main/videos/think_connect4.mp4" controls autoplay loop muted playsinline width="100%"></video>
**Wordle.** Which word still fits all the colour feedback? 52.0 % → 100 %.
<video src="https://huggingface.co/autotrust/GEV-26B-Decide/resolve/main/videos/think_wordle.mp4" controls autoplay loop muted playsinline width="100%"></video>
**Sudoku.** Which digit belongs in the highlighted cell? 76.0 % → 99.5 %.
<video src="https://huggingface.co/autotrust/GEV-26B-Decide/resolve/main/videos/think_sudoku.mp4" controls autoplay loop muted playsinline width="100%"></video>
One-move puzzles, 200 generated puzzles per game (threshold 0.8, budget 8,192 thinking tokens; text input, chess with the
| game | question (options) | chance | System 1 | **adaptive** | puzzles that thought |
|---|---|---:|---:|---:|---:|
| Minesweeper | which hidden cell is certainly safe (1 safe cell, 3 mines) | 25.0 | 21.0 | **86.0** | 100 % |
| Wordle | which word fits all the feedback (8 words) | 12.5 | 52.0 | **100.0** | 98 % |
| Connect Four | which column wins now or blocks (legal columns) | 14.5 | 54.0 | **99.5** | 95 % |
| Sudoku | which digit goes in the cell (1–9) | 11.1 | 76.0 | **99.5** | 92 % |
| 24 game | which expression equals 24 (6 expressions) | 16.7 | 89.0 | **100.0** | 74 % |
| Maze | first step towards the exit (2–4 directions) | 48.4 | 48.5 | **68.0** | 82 % |
| Chess (Lichess puzzles) | which move mates in one (16 moves, board image + FEN) | 6.3 | 46.5 | **78.0** | — |
Whole games and perception, where thinking helps little or not at all:
| check | System 1 | adaptive |
| Snake, one game (image + positions) | 3 food in 18 steps | 11 food in 90 steps (about 21 s per step) |
| Connect Four, 6 full games against a heuristic opponent | 1 win, 5 losses | 1 win, 4 losses, 1 draw |
| 2048, one game | 1,476 points | 1,016 points |
| Flappy Bird, one game | 0 pipes (23 frames) | 0 pipes (38 frames, about 35 s per frame) |
| Quick, Draw!, 320 real sketches, 16 answers (30 / 60 / 100 % of the strokes) | 46.9 / 73.1 / 94.4 % | 40.3 / 72.8 / 95.0 % |
Thinking pays off when the answer can be checked step by step against explicit rules: logic puzzles, tactics,
constraints, arithmetic. It does not help perception (sketches), reflexes (Flappy Bird) or long-horizon play (2048, full
Connect Four games), and every thought costs seconds. Details: `reports/thinking_games.json`.
1. **Fast distribution.** System 1 returns p1 in one pass.
2. **Think only when uncertain.** If the leading option of p1 is below the threshold (default 0.8), System 2 reasons in
Gemma-4's thinking mode over the same state, question and options. The reasoning length is yours to set, as with the
base model: `think_budget` caps the thinking tokens (default: no cap beyond the context window; the evaluations below
3. **Fold the reasoning in.** When the thinking channel closes, the answer-letter distribution p2 is read in one step,
and the result is p = ½ p1 + ½ p2. On its own, p2 is close to one-hot and over-confident; the equal mix keeps the
reasoning's accuracy and System 1's calibration.
The threshold and the mix were chosen on 1,754 questions from six public sets that are not part of the Decision Index
(test or validation splits, 300 random questions each; AQuA-RAT has 254):
| set | System 1 | **adaptive** | thinking on | always think |
|---|---:|---:|---:|---:|
| AQuA-RAT (math word problems) | 68.1 | **89.0** | 53.9 % | 90.2 |
| LogiQA (logical reasoning) | 55.3 | **80.3** | 63.7 % | 82.0 |
| StrategyQA (multi-hop yes/no) | 68.0 | **78.7** | 71.0 % | 79.0 |
| MedMCQA (medical) | 66.7 | **71.7** | 52.3 % | 73.7 |
| OpenBookQA (science) | 94.3 | **96.3** | 14.7 % | 95.7 |
| CommonsenseQA | 86.7 | **85.0** | 32.0 % | 83.3 |
| **all 1,754** | 73.3 | **83.4** | 47.8 % | 83.8 |