Two distinct GGUF releases trend on Hugging Face
SC117’s abliterated Qwen3.8 GGUF and Unsloth’s EmbeddingGemma 2 GGUF newly trended on Hugging Face.
The two reports cover different Hugging Face trending GGUF releases and do not contradict each other. On 2026-10-02, SC117’s Qwen3.8-Flash-Next-GSQ-RCO-abliterated weights were at rank 30: they keep ISTA-DASLab’s official GSQ-RCO quantization while replacing 144 residual-stream projection tensors across 48 layers with tensors from an orcarouter uncensored release, without recomputing learned GSQ values or scales. Available files are IQ3_S, IQ3_XXS, IQ2_XS, and Q2_0, about 38.0 GB to 55.1 GB, with a 262K context window under Apache-2.0, and the card says IQ3_S was verified tensor by tensor with BLAKE2b. On 2026-10-06, Unsloth’s Dynamic 3.0 GGUF of Google DeepMind’s EmbeddingGemma 2 ranked 29: a 740M-parameter Apache 2.0 model that places text (including code), images, video, and audio in one 768-dimensional space, using an about 270M text stack with optional 170M vision and 300M audio encoders. Context is 8,192 tokens across 100+ languages, with Matryoshka truncation to 128, 256, or 512 dimensions for up to 6x less storage; reported full-precision means are multilingual MTEB 61.36 and code MTEB 78.68, compared with 68.76 for EmbeddingGemma 1 on code MTEB.
- On 2026-10-02, SC117’s Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF was trending at #30 on Hugging Face.
- It keeps ISTA-DASLab’s official GSQ-RCO quantization and swaps 144 residual-stream projection tensors across all 48 layers from an orcarouter uncensored release, without recomputing GSQ values or scales.
- Listed quants are IQ3_S, IQ3_XXS, IQ2_XS, and Q2_0, about 38.0 GB (Q2_0) to 55.1 GB (IQ3_S), with a 262K context window under Apache-2.0; the IQ3_S file was checked per tensor with BLAKE2b.
- On 2026-10-06, Unsloth’s embeddinggemma-2-GGUF (Dynamic 3.0) was trending at #29.
- Google DeepMind’s EmbeddingGemma 2 is a 740M-parameter Apache 2.0 multimodal embedder: about a 270M text stack plus optional 170M vision and 300M audio encoders, mapping text (including code), images, video, and audio to 768 dimensions…
- Matryoshka truncation to 128, 256, or 512 dimensions (from 768) can cut vector storage up to 6x; cited full-precision scores are multilingual MTEB 61.36 and code MTEB 78.68, versus EmbeddingGemma 1’s code MTEB 68.76.
Coverage timelineoldest first · each row is one article
- · 6d agoSC117/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF — new model trending #30 on Hugging Face
Hugging Face trending models· 34
SC117 published abliterated GGUF quants of Qwen3.8-Flash-Next that swap 144 tensors while keeping GSQ-RCO scales.
- · 2d agounsloth/embeddinggemma-2-GGUF — new model trending #29 on Hugging Face
Hugging Face trending models· 52
Unsloth released GGUF weights for Google DeepMind's EmbeddingGemma 2, a 740M multimodal embedding model.