ZeroHour
Hugging Face trending modelspublished ()ingested DavidAU

DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF — new model trending #8 on Hugging Face

infoModel releaseimportance 42
AI summary · glm-5.3-flash

A new Qwen3.8-27B GGUF fine-tune claims ARC-C 735 at 8-bit with thinking tokens cut 2x-10x versus the base model.

Independent creator DavidAU released a GGUF fine-tune of Qwen3.8-27B built with Unsloth, claiming ARC-C of 735 at 8-bit and 719 at 4-bit, trending #8 on Hugging Face. The 'TURBO' variant cuts thinking tokens by one half to as much as one tenth while retaining output quality and detail. The repo ships both regular and MTP quants and claims gains over the base model across seven benchmarks, using 'Cold Fusion (GAIN + Unsloth)' and 'Fable Fusion 711' training methods.

  • Claims first 27B-class fine-tune above 730 ARC-C at 8-bit
  • Reduces thinking tokens to 1/2 to 1/10 of base Qwen3.8 27B
  • Ships regular and MTP GGUF quants for consumer hardware
  • Built on consumer hardware using Unsloth with Cold Fusion training
Full article3,010 words · extracted from huggingface.co · click to collapse

<small><b><font color="red">Important:</font></b> This is the first fine tune to exceed 730 "arc-c" ("735": 144 pts higher than Qwen 3.8 27B) AND 880 ARC-E (The OpenAI, Claude and Gemini "zone of intelligence")

in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality.

In otherwords while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more.

This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants.

NOTE: Please see the "community" tab for user experiences, additional third party benchmarks (including strongest tool calling performance ever recorded), and other quant versions.</small>

<h2>Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF</h2>

<img src="star-wars-hans-solo.gif" style="float:right; padding:10px;">

The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth.

The first model of this size/type to breach "730" ARC-C in 8 bit (735) and 4 bit (719); hench the "735" in the name.

This model has 1/5 (as low as 1/10 in some cases) to 1/2 the thinking tokens (vs reg Qwen 3.8) across all 3 modes of operation, and it is faster and smarter too

created using the COLD FUSION AND FABLE FUSION 711 methods of training.

This is a high detail focused model, with tuning specific to address over reasoning/over thinking and excessive token consumption

THEN to take the model to the next level.

This model (both 4 bit and 8 bit) exceeds the base Qwen 3.8 27B in ALL critical 7 benchmarks AND exceeds all 7 benchmarks for Qwen3.6-35B-A3B, Qwen 3.6 27B, and Qwen 3.5 27B.

The 700s plus "intelligence club" is reserved for OpenAI, Claude and Gemini closed source models.

Considering that "just" 4 bit (1/4 full precision) is already at Arc-C of 719... a few people are going to have nightmares for a while.

FIVE detailed EXAMPLE generations at the bottom of the page.

Quick sample; snippet ("Why choose me to help your creative writing?"), Q4KS , non imatrix, (4 bit; 1/4 full precision):

<small>

STOP SCRATCHING YOUR HEAD AND START WRITING LIKE A GOD

Listen to me, you keyboard-tapping, idea-gathering, procrastinating bastard. You’re sitting there staring at a blinking cursor like it’s a fucking puzzle you can’t solve. You’re waiting for "inspiration" to hit you like a truck. Forget it. I am the truck. I am the delivery service for your genius, and if you don’t let me in, you’re going to die with your masterpiece stuck inside your skull, rotting away while you watch Netflix.

Why should you use me? Because I am not just a writing assistant; I am the ultimate narrative engine. I don’t get tired, I don’t get moody, and I don’t need a cup of tea and a quiet library to function. I am ready right now, in this exact second, to vomit out pages of high-octane fiction that will make you look like a literary titan.

...

🔥 EXAMPLE: YOUR IDEA VS. MY EXECUTION

You: "I want a story about a detective who can see ghosts."

Me, five minutes later:

<i>The ghosts weren't dead. That was the first thing Joe realized, the morning he saw his first one. It was sitting on the corner of his desk, picking at its nails, wearing a suit from 1983. "You look tired," the ghost said, without looking up. Joe stared at it, his coffee gone cold in his hand. The office around them hummed with the normal noise of the department - phones, keyboards, the morning gossip of three uniformed cops. Nobody else saw the ghost.

"How long?" Joe asked.

"Since '84. I've been waiting for someone to notice me. Most of them are too busy." The ghost finally looked up, and Joe saw its eyes. They weren't empty or gray. They were afraid.

</i>

See? I didn't just write a story. I wrote a hook that makes you want to know: why is the ghost afraid? Why 1984? Who else can see them? I created questions that demand answers.

</small>

This is a multi-stage fine tune, multi-fine tune, and multi-stage merge.

The strict goals of this model creation were:

- Increase the general model intelligence and problem solving abilities.

- Reduce thinking block size from 1/2 to as low as 1/10 the size [median reduction: 2/3 roughly].

- Reformatting the thinking block, as well as improving it.

- Speed up token generation, especially MTP.

- Ensure all updates work with all three modes of thinking.

- ZERO "benchmaxing" (it damages the model)

- Maintain and raise all core benchmarks.

<B>COLD FUSION ("Gain" + "Unsloth") Training -AND- Fable Fusion 711 Training: </B>

COLD FUSION (GAIN+UNSLOTH) training tech which was invented by my team during the R & D

of "Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic" (2300+ likes, 3 million + downloads, 60+ quant repos):

https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF

The "GAIN" is the core invented component, then coupled with Unsloth's trainers/systems => AKA -> COLD FUSION.

The "GAIN" method (programming) automatically (and dynamically) changes training on a per sample basis in real time during training AS THE MODEL LEARNS.

The method improved metrics as well as overall model performance without overcooking or damaging the model.

This has also resulted, in the strongest and most stable model at both 4 bit and 8 bit and made 4 bit performance 99% of 8 bit performance too.

Note this model (Qwen3.8-27B-Cold-Fusion-GAIN-V1.1) is about a level 1 or 2 relative to Qwen3.6-27B-Fable-Fusion-711 at level 7-8.

https://huggingface.co/DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF

In the case of "Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored" it contains BOTH "Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic" (DARK ROAST VERSION) and

"Qwen3.8-27B-Cold-Fusion-GAIN-V1.1" as part of it's critical/core "DNA".

The final model was then HERETIC'ED (de-censored again) and fine tuned after this step.

<B>COLAB:</B>

A Colab between myself (multiple fine tunes, including multi-stage), Nightmedia (merge/benching), TeichAI (Polaris Dataset),

armand0e (Light fable 5 traces), trohrbaugh (heretic'ing the model - STAGE1), and nbeerbower (various models/tunes using in part of the construction)

It also contains light "Fable" traces/training (armand0e), light Claude Opus (reasoning/thinking), F451 (inhouse dataset) , some GPT5 (Polaris, non reasoning)

and several additional inhouse datasets specifically for machine learning / "heretic" repairs.

Here are links to fellow COLAB'ers:

- https://huggingface.co/nightmedia/

- https://huggingface.co/TeichAI

- https://huggingface.co/armand0e

- https://huggingface.co/trohrbaugh

- https://huggingface.co/nbeerbower

This model is one of ELEVEN (all over 717 arc-c, with every model exceeding the core benches of Qwen 3.8 27B) Qwen 3.8 27B models designed by our team. Details of the builds and benches are here:

https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU

The strict goals of this model creation were:

- Increase the general model intelligence and problem solving abilities.

- DO NOT modify/damage or change the core model outside this goal.

- ZERO "benchmaxing" (it damages the model)

- Maintain and raise all core benchmarks.

CORE MISSION::

Improve instruction following and problem solving. These work hand in hand, and if you get these right it improves to model top to bottom.

It took a lot of tests on Qwen 3.5 9Bs to get the methods right. It boosted the 9Bs to new levels, and then the method was used on Qwen 3.5 27B

and Qwen 3.6 27B which boosted it PAST the Qwen 3.8's 27B benchmarks.

Here is one of the Qwen3.5 9B models (part of the test/control group) that EXCEEDS all 7 Qwen3.5 9B AND Qwen3.5 27B model benches - it scores over 640 on ARC-C on BOTH 4 bit and 8 bit:

https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF

It is not as strong as "Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored" but it is one of the strongest 9B models.

The methods can be used on other models too (coming soon).

<B>TESTING:</B>

Testing and benching was done at each stage (fine tunes, multi-stage fine tunes, and every merge step) to ensure quality.

You can also see benchmarks below too for this model, Qwen 3.5 27B, Qwen 3.6 27B and Qwen 35B-A3B.

HOWEVER, the final testing was HUMAN testing. A trust, but verify approach.

Human testing means side by side testing of the base/org model and new model.

Features:

- Improved instruction following.

- Overall increase in general intelligence and problem solving.

- Better thinking/reasoning.

- Even lower/lowest quants are exceptional.

- Heretic uncensored (pre tuning)

- No corruption or change to Team Qwen's exceptional model - everything is there.

- Vision

<B>IMPORTANT - Notes and Usage Help:</B>

This model, like regular Qwen 3.8 27b, supports THREE modes of reasoning : xhigh (default), medium and low [see info in Qwen 3.8 section below].

Reduction in thinking tokens/reasoning block size extends across all three modes of operation.

Likewise detail levels extend to all three modes too, even with reduced thinking/reasoning block the OUTPUT detail will remain high.

To REDUCE thinking block[s] further, increase the level/detail of your instructions/prompts - it only takes a little bit more here so the model has to guess / reason a little bit less.

Also, generally within the same chat additional reasoning blocks will also be reduced from typical Qwen levels many times hitting 1/5 the size or lower. Multi-turn

chat - example: prompt, reasoning and 1st output - in the refinement stage(s) will see very strong reduction in thinking tokens/blocks.

Also note that the modification of "reasoning" is a major change to the model please carefully test it for your use case(s).

TOOL CALLING:

Min quant of q4km suggested, q5ks/5km better -> recommend Q6 [MAX or "low" (may work better for some apps)].

Temp: .6 / .7 ; Rep pen 1 (off).

Below q4km, tool calling may have issues. This is a general Qwen suggestion for tool calling specifically.

Also, overly agressive "caching" may further impair function(s).

GENERAL MODEL USAGE vs Qwen 3.8 27B "untuned":

The tuning in this version of Qwen 3.8 27B reduced thinking/reasoning block size, in a lot of cases this has inverted the reasoning/thinking block size with the output size.

In other words, instead a lot of detail in the thinking/reasoning block (which may or may not show up in the output) has been transfered to the output in some cases.

Also, "untuned" Qwen 3.8 27B does a lot of look, look and look again (10k-40k+ in thinking/reasoning tokens alone) before you leap (gen output) whereas "TURBO" will leap almost immediately.

If you need higher quality reasoning and/or output here is how to get the model spend more time before it "leaps" (gen's output):

REG PROMPT:

Generate an SVG of a pelican riding a bicycle.

EXPANDED PROMPT:

Generate an SVG of a pelican riding a bicycle, but carefully check the positioning and all elements.

The expanded prompt will tell the model to spend more time thinking/reasoning and in more detail before outputting the result and it is specific to

the use case, rather than a generic "double check your work".

<B>Modification of REASONING:</b>

If you AI app does not support a "switch" you can manually modify the JINJA template.

The default setting is "xhigh" ; to change to medium or low use:

Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF