autotrust published JEV-27B-VL, a vision model that makes calibrated System 1 decisions over images in one pass.
autotrust/JEV-27B-VL adds vision to JEV-27B. System 2 is the unmodified Qwen3.8-27B with optional step-by-step reasoning on images. System 1 answers yes/no, multi-choice, or 0–5 ratings over text and images and returns a calibrated probability for every option in one forward pass, with prompts up to 256K tokens. Reported demos completed 75% of 20 robot-arm pick-and-place scenes and 95% of 60 multi-step computer-use tasks, and cover-only ranking reached 0.727 AUC on MicroLens-100k.
System 1 returns calibrated probabilities for up to 256 options in one pass.
Robot-arm pick-and-place succeeded in 75% of 20 MuJoCo scenes.
Computer-use loop finished 95% of 60 multi-step browser tasks.
Zero-shot cover ranking matched collaborative-filtering AUC on MicroLens-100k.
Hidden-sentence checks stayed correct through about 250K tokens of context.
Full article2,909 words · extracted from huggingface.co · click to collapse
# autotrust/JEV-27B-VL
### JEV-27B that can see: the same System 1 decisions and System 2, now on images
**autotrust/JEV-27B-VL** is [autotrust/JEV-27B](https://huggingface.co/autotrust/JEV-27B) with vision.
| | what it does | output |
|---|---|---|
| **System 1** (`POST /v1/decide`) | typed decisions: yes/no · pick one of 2–256 options · rate 0–5, over text and images; prompts up to 256K tokens | a calibrated probability for every option, in one forward pass |
| **System 2** (`autotrust/JEV-27B-VL`) | the unmodified Qwen3.8-27B, optionally thinking step by step, with image input | text / reasoning |
## New (3 October 2026): robot arm and computer use
Every step below is **one System 1 decision**: a camera image or a screenshot in, a probability for every action out, in
a single forward pass.
**Robot arm: pick and place from a camera image.** At every step System 1 looks at the top camera image and answers two
questions: is the target left or right of the gripper, and above or below it? The arm moves accordingly and halves its
step whenever an answer flips. It grasps the cube, carries it and drops it in the tray (MuJoCo simulation).
* **Matches collaborative filtering with zero behaviour data:** the same AUC as item-based CF learned from 59,045 users'
watch histories (difference 0.000, 95% interval −0.041 to +0.040), and a higher top-5 hit rate (59% vs 49%).
* **Solves the cold-start problem:** collaborative filtering needs co-watch history; JEV only needs the cover, so new videos
and new creators can be recommended from the first second.
* **Pictures beat words:** covers alone score +0.078 AUC over titles (95% interval +0.031 to +0.126).
* **Fast enough for a ranking stage:** 20 candidate covers scored in about 2.6 s on one GPU in the live demo.
Code, data pipeline and a live web demo: [JEV-27B-DEMO / 07-video-recommendation](https://github.com/yuhai-china/JEV-27B-DEMO/tree/master/07-video-recommendation).
## Watch it decide
Every move below is one System 1 decision: a probability for each possible action, read off in a single forward pass. The
panels show the action probabilities and the time each decision took.
**Super Mario Bros.** Each step is a movement choice plus a jump decision, about 150 ms per decision (video at 5× speed).