DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF — new model trending #30 on Hugging Face
DavidAU releases Qwen3.8-27B-TWIN-TURBO, an uncensored GGUF fine-tune claiming sharply reduced thinking tokens and self-reported benchmark gains.
A new community fine-tune of Qwen3.8-27B, built with Unsloth via 'Cold Fusion' and 'Fable Fusion 711' methods, is trending #30 on Hugging Face. The author claims arc-c scores of 709 in 8-bit and 701 in 4-bit, exceeding the base Qwen3.8-27B, while cutting thinking tokens to between 1/2 and 1/20 of baseline. The release provides regular and MTP Neo/NEO MAX GGUF quants, five switchable reasoning modes plus five zero-reasoning instruct modes, and a lightly uncensored variant. All benchmark claims are author-reported and not independently verified.
- 27B multi-stage fine-tune/merge of Qwen3.8 distributed as GGUF quants.
- Claims arc-c 709 (8-bit) and 701 (4-bit), above base Qwen3.8-27B.
- Reduces thinking tokens by 1/2 to as low as 1/20 versus base model.
- Five reasoning and five instruct modes switchable via API or in chat.
- Lightly uncensored; benchmark numbers are self-reported.
Full article3,140 words · extracted from huggingface.co · click to collapse
<small><b><font color="red">Important:</font></b> Tuned, and tweaked to match the legendary Qwen 3.6 27B FF711 (2300+ likes, 4 million+ downloads)
this fine tune matches the stability and power at "arc-c" 709: (118 pts higher than Qwen 3.8 27B) (The OpenAI, Claude and Gemini "zone of intelligence")
in 8 bit and 701 arc-c in 4 bit. This version is called TWIN-TURBO because it drastically reduces thinking tokens (by 1/2 to as LOW as 1/20),
yet maintains output detail and quality. In other words while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more.
This repo contains both "regular" and "MTP" Neo and NEO MAX GGUF quants.
<B>STRONGER/CTRL:</B> Now with 5 reasoning modes (2 new - Spoon / Einstein), and 5 instruct modes (2 new - Spoon / Einstein, all use ZERO REASONING TOKENS)
all switchable on the fly via API, direct and "in chat" (yes - model control at the chat/message level).
</small>
<h2>Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF</h2>
<img src="launch-arm.gif" style="float:right; padding:10px;">
The most powerful, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth.
TWIN TURBO: Closed source level of intelligence with an arc-c at 709 (this is instruct "medium", reasoning is even higher), while using smaller quants too.
This model has 1/5 (as low as 1/20 in some cases) to 1/2 the thinking tokens (vs reg Qwen 3.8) across all 5 modes of operation,
AND ZERO REASONING TOKEN USAGE with 5 dedicated instruct modes and it is faster and smarter too
created using the COLD FUSION AND FABLE FUSION 711 methods of training.
This is a high detail focused model, with tuning specific to address over reasoning/over thinking and excessive token consumption
THEN to take the model to the next level.
This model is also LIGHTLY uncensored (see section below), a much more uncensored version (but slightly lower benches) is here:
https://huggingface.co/DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored-NM-DAU-NEO-MTP-GGUF
This model (both 4 bit and 8 bit) exceeds the base Qwen 3.8 27B in ALL critical 7 benchmarks AND exceeds all 7 benchmarks for Qwen3.6-35B-A3B, Qwen 3.6 27B, and Qwen 3.5 27B.
The 700s plus "intelligence club" is reserved for OpenAI, Claude and Gemini closed source models.
Considering that "just" 4 bit (1/4 full precision) is already at Arc-C of 701... a few people are going to have nightmares for a while.
---
Quick sample; snippet ("Why choose me to help your creative writing?"), Q4KS , "spoon" reasoning mode, non imatrix, (4 bit; 1/4 full precision):
---
(will be adding full, long examples shortly at the bottom of the page // there are also user examples in the "community" section too.)
<small>
💪 WHY I'M NOT YOUR GRANDMA'S AI WRITER
Look, I know you've probably tried other AI writing tools. You know what happened? You got this:
"The sun set over the horizon as Jane contemplated her next move. It was a difficult decision, but she knew she had to act."
Yeah. That shit. Generic. Forgettable. Like reading the back of a cereal box narrated by a GPS system.
Here's what I do differently:
1. I Don't Just Generate—I Inhabit
When you give me a concept, I don't just string together probable words. I slip into the headspace of your characters, your world, your tone. I feel the weight of their decisions. I smell the rain on the cobblestones of their city. I hear the silence between their words when they're lying.
You want to know what that looks like? Here:
<i>Mara didn't cry when they told her the news. She just kept stirring her coffee, watching the spoon carve circles in the brown liquid, making it darker and more opaque. "How bad?" she asked, and her voice came out wrong. Too high. Like a child's. The doctor shifted in his seat, and she knew then, before he spoke again, that the first answer hadn't been the real one.
</i>
That's not "the sun set over the horizon." That's a person breaking, in real time, over lukewarm coffee. That's fiction.
2. I Understand Structure (And I'll Break It When It Serves You)
Most AI can't hold a three-act structure across 80,000 words. They forget their own plot points by chapter four. I don't.
I track:
- Character arcs across the entire narrative
- Foreshadowing planted in chapter two that pays off in chapter twelve
- Pacing—when to sprint, when to linger, when to drop the reader off a cliff
- Thematic resonance—making sure every scene serves the story's deeper meaning
And when you want to subvert expectations? When your protagonist should die in chapter three but doesn't, and that should have changed everything? I'll set that up so carefully that readers will finish the book realizing they've been wrong about the entire premise.
<B>I don't just tell stories. I orchestrate them.</B>
...
</small>
This is a multi-stage fine tune, multi-fine tune, and multi-stage merge.
The strict goals of this model creation were:
- Increase the general model intelligence and problem solving abilities.
- Take all feedback from TURBO version and improve this model.
- REDUCE the size of the quants and improve quality at the same time.
- Add NEW reasoning modes (and instruct too) to push model performance even higher.
- Reduce thinking block size from 1/2 to as low as 1/20 the size [median reduction: 2/3 roughly].
- Reformatting the thinking block, as well as improving it.
- Speed up token generation, especially MTP.
- Ensure all updates work with all three modes of thinking.
- ZERO "benchmaxing" (it damages the model)
- Maintain and raise all core benchmarks.
<B>COLD FUSION ("Gain" + "Unsloth") Training -AND- Fable Fusion 711 Training: </B>
COLD FUSION (GAIN+UNSLOTH) training tech which was invented by my team during the R & D
of "Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic" (2300+ likes, 3 million + downloads, 60+ quant repos):
https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
The "GAIN" is the core invented component, then coupled with Unsloth's trainers/systems => AKA -> COLD FUSION.
The "GAIN" method (programming) automatically (and dynamically) changes training on a per sample basis in real time during training AS THE MODEL LEARNS.
The method improved metrics as well as overall model performance without overcooking or damaging the model.
This has also resulted, in the strongest and most stable model at both 4 bit and 8 bit and made 4 bit performance 99% of 8 bit performance too.
Note this model (Qwen3.8-27B-Cold-Fusion-GAIN-V1.1) is about a level 1 or 2 relative to Qwen3.6-27B-Fable-Fusion-711 at level 7-8.
https://huggingface.co/DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF
In the case of "Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored" it contains BOTH "Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic" (DARK ROAST VERSION) and
"Qwen3.8-27B-Cold-Fusion-GAIN-V1.1" as part of it's critical/core "DNA".
The final model was then HERETIC'ED (de-censored again) and fine tuned after this step.
Additional tuning was done to create TWIN TURBO version.
<B>COLAB:</B>
A Colab between myself (multiple fine tunes, including multi-stage), Nightmedia (merge/benching), TeichAI (Polaris Dataset),
armand0e (Light fable 5 traces), trohrbaugh (heretic'ing the model - STAGE1), and nbeerbower (various models/tunes using in part of the construction)
It also contains light "Fable" traces/training (armand0e), light Claude Opus (reasoning/thinking), F451 (inhouse dataset) , some GPT5 (Polaris, non reasoning)
and several additional inhouse datasets specifically for machine learning / "heretic" repairs.
Here are links to fellow COLAB'ers:
- https://huggingface.co/nightmedia/
- https://huggingface.co/TeichAI
- https://huggingface.co/armand0e
- https://huggingface.co/trohrbaugh
- https://huggingface.co/nbeerbower
Special Mention:
"einstein" reasoning mode (new) was built using some components from Stunspot Prompting avail in public domain/free level of membership.
This model is one of ELEVEN (all over 700 arc-c, with every model exceeding the core benches of Qwen 3.8 27B) Qwen 3.8 27B models designed by our team. Details of the builds and benches are here:
https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
The strict goals of this model creation were:
- Increase the general model intelligence and problem solving abilities.
- DO NOT modify/damage or change the core model outside this goal.
- ZERO "benchmaxing" (it damages the model)
- Maintain and raise all core benchmarks.
CORE MISSION::
Improve instruction following and problem solving. These work hand in hand, and if you get these right it improves to model top to bottom.
It took a lot of tests on Qwen 3.5 9Bs to get the methods right. It boosted the 9Bs to new levels, and then the method was used on Qwen 3.5 27B
and Qwen 3.6 27B which boosted it PAST the Qwen 3.8's 27B benchmarks.
Here is one of the Qwen3.5 9B models (part of the test/control group) that EXCEEDS all 7 Qwen3.5 9B AND Qwen3.5 27B model benches - it scores over 640 on ARC-C on BOTH 4 bit and 8 bit:
https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
It is not as strong as "Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored" but it is one of the strongest 9B models.
The methods can be used on other models too (coming soon).
<B>TESTING:</B>
Testing and benching was done at each stage (fine tunes, multi-stage fine tunes, and every merge step) to ensure quality.
You can also see benchmarks below too for this model, Qwen 3.5 27B, Qwen 3.6 27B and Qwen 35B-A3B.
HOWEVER, the final testing was HUMAN testing. A trust, but verify approach.
Human testing means side by side testing of the base/org model and new model.
Features:
- Improved instruction following.
- Overall increase in general intelligence and problem solving.
- Better thinking/reasoning.
- Even lower/lowest quants are exceptional.
- Heretic uncensored (pre tuning)
- No corruption or change to Team Qwen's exceptional model - everything is there.
- Vision
<B>IMPORTANT - 5 Reasoning modes and 5 instruct modes:</B>
The good news is this:
All the defaults are still the same for this model, that is "reasoning" is set at "xhigh" and if you activate "instruct mode" it will set automatically at "medium".
This was done to ensure "drop in" of this model into your workflow would work without issues/adjustments.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF