nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2 — new model trending #30 on Hugging Face · ZeroHour
nerkyor released a Qwen3.8-27B post-train that cuts 94K truncations and lifts GPQA, MMLU, and coding scores.
Hugging Face user nerkyor published Qwen3.8-27B-Coder390, a multi-round SFT and RLOO post-train of official Qwen3.8-27B meant to stop unproductive re-reasoning after the model already has an answer. Static FP8 scores are GPQA 178/198, MMLU 450/500, and LiveCodeBench 90/100, versus 177/444/83 for the original FP8 checkpoint. 94K truncations fall from 4 to 1 on GPQA and from 13 to 3 on LCB. The repo includes BF16, FP8, NVFP4, INT8, INT4, and GGUF builds, plus MTP and DFlash2 speculative decoding at about 91 and 180 tokens/s versus 46 without.
Alternating SFT and RLOO on a Qwen3.8-27B EfficientThink base.
A model post-trained with several rounds of SFT and RLOO on top of [Qwen3.8-27B-EfficientThink-SFT-SimPO-DFlash2](https://huggingface.co/nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2), which is itself the official [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) post-trained with SFT and SimPO. It is built to fix the original model's habit of failing to stop: it already has an answer, yet keeps re-deriving until it hits the 94K cap with no final answer. This repository provides five safetensors tiers, **BF16**, **static FP8 Block128**, **NVFP4 W4A16**, **NVFP4 W4A4**, and **NVFP4 W4A4-W8A8** (mixed precision), all fully multimodal (text + image/video) with the official BF16 MTP bundled; all load in SGLang without patches, and the FP8 package is also smoke-tested on vLLM. It also has **NInfer W4A4 / W4A4-W8A8** (`NVFP4-NInfer/`), **INT8 W8A8** (`INT8/W8A8/`), **INT8 W8A16** (`INT8/W8A16/`, text + image), **INT4 W4A16** (`INT4/W4A16/`, text + image, vLLM), and **GGUF Q2 / Q3 / Q4 LynnStyle (plus Q8 MTP versions), Q6_K, and Q8_0** (`GGUF/`, `GGUF-NInfer/`, built-in MTP).
- **Far fewer 94K truncations**: under the same protocol, 4 → 1 on GPQA and 13 → 3 on LCB.
- **No regression on any suite**: GPQA **178/198**, MMLU **450/500**, LCB **90/100** (static FP8), vs. 177 / 444 / 83 for the original FP8.
- **Lossless quantization**: GPQA and LCB are identical across BF16, dynamic FP8, and static FP8.
- **Two speculative-decoding options (mutually exclusive)**: about 180 tok/s per request with DFlash2 and 91 tok/s with MTP, vs. about 46 tok/s without speculation.
**Lineage**: original Qwen3.8 27B (FP8; 177 / 444 / 83) → our SFT + SimPO110 final, i.e. the [EfficientThink base](https://huggingface.co/nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2) (171 / 442 / 89) → K3 continuation SFT → week-2 SFT (week2dose, update-225) → first RLOO round (182 groups) → another SFT round (merge-sft-100), giving sft-base-rloo (177 / 448 / 90) → second RLOO round (172 groups) → **Coder390** (178 / 445 / 90, dynamic FP8). Numbers are GPQA / MMLU / LCB, all full suites under the same 100K protocol.
**Training base**: this model starts from our own EfficientThink SFT + SimPO model (the SimPO110 final) and goes through several rounds of SFT and RLOO, which is itself a post-trained version of the official Qwen3.8-27B: [Qwen3.8-27B-EfficientThink-SFT-SimPO-DFlash2](https://huggingface.co/nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2).
| Name part | Meaning |
|---|---|
| Coder | Strengthened coding |
| 390 | GPQA, MMLU, and LCB all reach 90 |
| EfficientThink | Training goal: remove unproductive reasoning tails while keeping necessary long reasoning |
| Opus5.5, GPT6Astra | Wrote the gold answers and teacher trajectories |
| Grok4.7 | Host of this training run; checked and filtered data throughout |
| DSV4Pro, K3 | Their trajectories served as teacher-model data; K3 also performed the RLOO value review and wrote part of the problems |
| SFT-RLOO | Method: alternating SFT and RLOO rounds on an SFT (incl. SimPO) base |
**The problem**: after the model already has a local answer, it keeps re-running the same derivation with Wait / Actually until the context is full, and the final answer is empty. In most of these cases the model can solve the problem; it just does not stop.
**RLOO data**: 8 trajectories sampled per problem; only groups with both correct and wrong trajectories are kept (all-correct groups are dropped). Groups with 0 or 1 correct trajectory receive one reviewed short teacher trajectory, which replaces the shortest wrong trajectory in that group (27 groups). Final set: 172 groups, 1,376 trajectories: 859 correct and 517 wrong, including 68 empty answers.
**Reward**: penalties apply only to wrong trajectories; long correct reasoning still earns a positive reward, so necessary long reasoning is not suppressed.
| Case | Reward |
|---|---:|
| Short and correct (<24K) | +1.05 |
| Long and correct | +1.0 |
| Short and wrong (<24K) | −0.2 |
| Wrong, 24K–48K | −0.5 |
| Wrong, ≥48K | −0.7 |
| Reached 94K with an answer letter, but wrong | −0.9 |
| Empty answer | −1.0 |
**Result**: under the same protocol, 94K truncations drop from 4 to 1 on GPQA and from 13 to 3 on LCB (FP8) compared with the original Qwen3.8 27B (FP8), and none of the three scores falls below it (all FP8 figures; see the NVFP4 scores below).
## Scores
**Protocol**: the same Coder390 merged weights; two RTX PRO 6000 GPUs, C8 per GPU, 100K context, 94,208-token generation cap, no client-side timeout, DFlash2 draft, SGLang; full GPQA 198, MMLU 500, and LCB 100. Empty answers count as wrong, truncated-but-correct answers count as correct, and every failed sample stays in the denominator.
| Suite | Original Qwen3.8 27B (FP8) | BF16 | Dynamic FP8 | Static FP8 |
GPQA and LCB are identical across all three precisions. MMLU varies between 445 and 450, which is sampling noise rather than a precision effect. Dynamic FP8 means online FP8 quantization of the BF16 weights at inference time; it is not shipped as a separate package.
### All tiers: scores, reasoning length, 94K truncation, and empty answers
Reasoning length is `usage.reasoning_tokens`; a 94K truncation means the output hit the generation cap. LCB empty answers are the number of problems for which no runnable code could be extracted. The original, BF16, dynamic FP8, and static FP8 rows use two RTX PRO 6000 GPUs with C8 per GPU; the three NVFP4 tiers use a single RTX PRO 6000 at C8; the rest of the protocol is the same.
Bold marks values better than the original Qwen3.8 27B (FP8): a higher score, or lower P50 / P70 / P90, 94K truncations, or empty answers; values equal to or worse than the original are not bolded.
LCB improves the most: all 13 of the original's 94K truncations ended with empty code (13 truncations, 13 empty answers), and the Coder390 tiers cut this to 2–4. The 3 for W4A4 are problems 0, 53, and 96 (problem ids in the score file), likewise truncated with no code; of the 4 for W4A16, problems 6, 7, and 54 have no code, and problem 53 has 20 characters of code but is wrong; of the 2 for W4A4-W8A8, problem 11 has no code, and problem 6 has 2,338 characters of code but is wrong. On GPQA the original had 4 truncations at the 94K cap and the tiers have 1–2; on MMLU the original had 1 and the tiers have 0–2, with no clear change.
<details open>
<summary>Details for six precisions (bucketed by reasoning length)</summary>
Buckets are "correct / questions in bucket", split by reasoning length, left-closed and right-open.
**Protocol**: one RTX PRO 6000, C8, 100K context, 94,208-token generation cap, no client-side timeout, DFlash2 draft, SGLang; full GPQA 198, MMLU 500, and LCB 100.
Reasoning length, 94K truncations, and empty answers for the three NVFP4 tiers are in the "All tiers" table above. LCB correctness is judged by each problem's `pass` field. Raw records: `evaluation/SCORES_NVFP4_W4A16.json`, `evaluation/SCORES_NVFP4_W4A4.json`, and `evaluation/SCORES_NVFP4_MIXED.json` (the W4A4-W8A8 tier).
**Protocol**: scores carried over from the official triad run on 2026-10-05 on the same NInfer text path; one RTX PRO 6000, official NInfer engine (`Neroued/ninfer`), C8, 100K context, 94,208 generation cap, client not timed, **no speculation**; full GPQA 198 / MMLU 500 / LCB 100; empty answers count as wrong, truncations that still answer correctly count as correct, and every failure stays in the denominator. The packages uploaded now add a Q8 MTP head, the DFlash2 draft, and the proposal head on top of that text path; the text path is unchanged and the full triad was not re-run. Only the NInfer engine can load `.ninfer`; SGLang / vLLM / llama.cpp / transformers cannot.
Reasoning length is `usage.reasoning_tokens` (i.e. `completion_tokens_details.reasoning_tokens`); a 94K truncation means the generation cap was hit. LCB empty answers are the number of problems for which no runnable code could be extracted. Bold means better than the original Qwen3.8 27B (FP8).
For NInfer W4A4, the 2 LCB empty answers are problems 11 and 15, both truncated with no code; the 2 GPQA empty answers are problems 79 and 127, both truncated with an empty pred. MMLU has no truncations and no empty answers. For NInfer W4A4-W8A8, the 3 LCB empty answers are problems 17, 53, and 54 (problem 17 stopped normally with no code; 53 and 54 were truncated with no code); the 2 GPQA empty answers are problems 79 and 127, both truncated with an empty pred. MMLU has no truncations and no empty answers.
<details open>
<summary>Details for the two NInfer tiers (by reasoning-length bucket)</summary>