autotrust/GEV-26B-Decide-NVFP4 — new model trending #30 on Hugging Face
autotrust published GEV-26B-Decide-NVFP4, an NVFP4 expert-quantized 26B model that fits in about 17 GB.
autotrust released GEV-26B-Decide-NVFP4, quantizing the 3,840 routed expert MLPs of GEV-26B-Decide to NVIDIA NVFP4 while leaving attention, routers, the vision tower, and decision heads in bf16. Download size falls from 49.5 GB to 18 GB, and vLLM weight memory from 51.1 GiB to 17.1 GiB. On System 1, GPQA Diamond moves from 43.9% to 44.9% and HLE text multiple choice from 8.4% to 9.7%, within run-to-run noise. The bf16 model scores 62.48 on Decision Index 0.2.1 with adaptive thinking.
48