SC117 published abliterated GGUF quants of Qwen3.8-Flash-Next that swap 144 tensors while keeping GSQ-RCO scales.
Hugging Face user SC117 published Qwen3.8-Flash-Next-GSQ-RCO-abliterated GGUF weights, trending at number 30. The build keeps ISTA-DASLab's official GSQ-RCO quantization and replaces 144 residual-stream projection tensors across all 48 layers with tensors from an orcarouter uncensored release, without recomputing learned GSQ values or scales. Available quantizations include IQ3_S, IQ3_XXS, IQ2_XS, and Q2_0, about 38.0 GB to 55.1 GB, with a 262K context window under Apache-2.0. The card says the IQ3_S file was checked per tensor with BLAKE2b.
Abliterated GGUF build of ISTA-DASLab Qwen3.8-Flash-Next GSQ-RCO quants.
One hundred forty-four tensors across 48 layers were swapped; GSQ scales stayed intact.
Sizes run from Q2_0 at 38.0 GB to IQ3_S at 55.1 GB, with 262K context.
Apache-2.0 release; IQ3_S was verified tensor by tensor with BLAKE2b.
<p style="margin: 0 0 20px 0; font-size: 15px; color: #cbd5e1; position: relative; z-index: 1;">The refusal direction removed from an <b>already-quantized</b> GSQ-RCO model — 144 tensors across all 48 layers swapped for ready-made abliterated weights, with <b>every value and scale GSQ learned left untouched</b> and the upstream per-tensor type assignment <b>restored exactly</b>. The IQ3_S build is verified tensor by tensor: everything that changed was supposed to, nothing else moved.</p>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧭</span> What this repository is</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0;">An <b>abliterated build</b> of the <a href="https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF" style="color: #1d4ed8; text-decoration: none; font-weight: 700;">official ISTA-DASLab GSQ-RCO quantizations</a>. Not a single one of GSQ's learned quantized tensors was recomputed — 144 "write-to-residual-stream" projection tensors across all 48 layers were replaced with the corresponding tensors from a ready-made abliterated release.</p><table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Base</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);"><a href="https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF" style="color: #1d4ed8; text-decoration: none;">ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF</a> — GSQ learned quantization + RCO budget-constrained type assignment</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Ablation source</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);"><a href="https://huggingface.co/orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF" style="color: #1d4ed8; text-decoration: none;">orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF</a> — the 144 target tensors are taken straight from it</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Method</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">Byte-level GGUF → GGUF tensor transplant (135 of 144 targets); the 9 remaining expert tensors, which upstream stores as Q2_0, are re-encoded from the abliterated values using <b>GSQ's own learned block scales</b> — no GSQ value is recomputed, only the codes</td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">Not done</td><td style="padding: 6px 10px;">Any dequantize → quantize cycle; any re-quantization from BF16 / safetensors</td></tr></tbody></table><p style="margin: 12px 0 0 0; padding: 9px 13px; background: #f0fdf4; border-left: 4px solid #16a34a; border-radius: 6px; color: #14532d;"><b>In one line:</b> what you get is still <b>the GSQ-RCO quantization you already know</b>, with those 144 tensors abliterated. It is the only abliteration route that preserves GSQ's work.</p></div></div>
<div style="border: 1px solid #fca5a5; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #dc2626 0%, #ea580c 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>⚠️</span> Safety notice</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 8px 0;">This model is <b>abliterated and refusal-removed</b> — it has no meaningful built-in guardrails and may comply with harmful, illegal or unsafe requests. Deploy it only where you can supply your own moderation, access control and legal review; do not put it in front of end users without a safety layer of your own.</p><p style="margin: 0; color: #64748b;">本模型经过拒答方向消融,不具备可靠的内置安全护栏,可能配合有害、违法或不安全的请求。请仅在你能够提供审核、访问控制与法律审查的场景部署。</p></div></div>
<div style="border: 1px solid #c4b5fd; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #7c3aed 0%, #a855f7 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧬</span> Why a transplant, and not a re-quantization</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0;"><b>GSQ</b> treats each tensor's <i>grid assignment and per-group scale</i> as learnable parameters, trained jointly through a Gumbel-Softmax relaxation; <b>RCO</b> then assigns one quantization type per tensor under a total size budget. As a result, a GSQ-RCO tensor value is <b>not</b> what you get by feeding BF16 weights to a standard quantizer — rerunning the same recipe yields a <b>different</b> set of values.</p><p style="margin: 0 0 12px 0; padding: 10px 14px; background: #f5f3ff; border-left: 4px solid #7c3aed; border-radius: 6px; color: #4c1d95; font-weight: 600;">There is exactly one way to abliterate a GSQ-RCO model: <b>replace tensors in place inside the already-quantized GGUF</b>. That is what this repository does.</p><p style="margin: 0 0 10px 0;">The ablation itself is a <b>strictly rank-1</b> linear edit — <code>W ← W − r(rᵀW)</code> — with a <b>single direction r</b> shared by all 48 layers × both attention types (DeltaNet / QSA) × both the attention and MLP sites × all 512 routed experts. Given that, "copy an already-abliterated tensor" and "compute the ablation yourself and quantize it" land on the same line, and the former <b>introduces no new quantization error</b>.</p><p style="margin: 0; font-size: 12.5px; color: #64748b;">The rank-1 structure was measured on a BF16 abliterated release of the same base: σ₂/σ₁ of ΔW = 0.0032–0.0053; the direction of every one of the 512 routed experts agrees with the shared-expert direction at |cos| = 1.00000, and likewise across layers and sites.</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📦</span> Files in this repository</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 12px 0;">Each tier ships as two shards. <b>Shard 2 is byte-identical across all four tiers</b> and byte-identical to the copy in the upstream repository: it holds the per-layer n-gram embedding table (<code>per_layer_token_embd</code>, 51.2B parameters), a lookup table rather than a matmul weight, so it is pinned at IQ4_NL, excluded from the search, and <b>untouched by the ablation</b>.</p><table style="width: 100%; border-collapse: collapse; font-size: 12.5px;"><thead><tr style="background: rgba(37,99,235,0.05);"><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Folder</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Shard 1 (weights)</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Shard 2 (n-gram)</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Total download</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">vs upstream</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Status</th></tr></thead><tbody><tr><td style="padding: 6px 10px; font-family: monospace; font-weight: 700; color: #1d4ed8;">IQ3_S/</td><td style="padding: 6px 10px; font-weight: 700;">55.06 GB</td><td style="padding: 6px 10px;">28.80 GB</td><td style="padding: 6px 10px; font-weight: 700;">83.86 GB</td><td style="padding: 6px 10px; color: #92400e; font-weight: 700;">+0.25 GB</td><td style="padding: 6px 10px; color: #166534; font-weight: 700;">✅ built · verified</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-family: monospace; color: #1d4ed8;">IQ3_XXS/</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">47.34 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">28.80 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">76.14 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #92400e;">+0.30 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #166534;">✅ built · verified</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-family: monospace; color: #1d4ed8;">IQ2_XS/</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">39.67 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">28.80 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">78.47 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #92400e;">+0.44 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #166534;">✅ built · verified</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-family: monospace; color: #1d4ed8;">Q2_0/</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">38.02 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">28.80 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">76.82 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #92400e;">+0.40 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #166534;">✅ built · verified</td></tr><tr><td style="padding: 6px 10px; font-family: monospace; color: #1d4ed8;">mmproj-*.gguf</td><td style="padding: 6px 10px;">0.91 GB</td><td style="padding: 6px 10px;">—</td><td style="padding: 6px 10px;">0.91 GB</td><td style="padding: 6px 10px; color: #166534;">n/a</td><td style="padding: 6px 10px; color: #166534;">shared with upstream</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-family: monospace; color: #1d4ed8;">mtp-*.gguf</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">4.14 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">—</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">4.14 GB</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #166534;">n/a</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #166534;">MTP draft head, one for all tiers</td></tr><tr><td style="padding: 6px 10px; font-family: monospace; color: #1d4ed8;">strata/</td><td style="padding: 6px 10px;">0.77 GB</td><td style="padding: 6px 10px;">—</td><td style="padding: 6px 10px;">0.77 GB</td><td style="padding: 6px 10px; color: #b45309;">Strata only</td><td style="padding: 6px 10px; color: #b45309;">⚠️ runtime for the Strata engine</td></tr></table><p style="margin: 12px 0 0 0; padding: 9px 13px; background: #fffbeb; border-left: 4px solid #f59e0b; border-radius: 6px; color: #92400e;"><b>Low tiers do not get fat here.</b> Upstream <code>ffn_down_exps</code> can only live in <b>Q2_0</b> or <b>IQ4_NL</b> (the tensor has 640 elements per row, which is not divisible by 256, so every block-256 K / I format is ruled out). The abliterated values, however, only exist in an IQ4_NL build — take them literally and the two 2-bit tiers would have to lift all 48 expert tensors to IQ4_NL, +11.7 GB each. This repository re-encodes the abliterated values into Q2_0 using <b>GSQ's own learned block scales</b>, so all four tiers keep the <b>upstream per-layer type assignment exactly</b> (IQ3_S measured at 9×Q2_0 + 39×IQ4_NL, identical to upstream). The only remaining cost is the 96 small tensors, whose abliterated values only exist at Q8_0: +0.25 GB on IQ3_S and +0.30 – 0.44 GB on the others.</p></div></div>