Qwen3.8-27B base powers two Apache-2.0 open-weight releases trending on Hugging Face: Prism ML's 1.72-bit ternary model and Altworld's writing-focused Hemmingway-1
Three Hugging Face trending reports in one week trace back to the same base model: Prism ML's Ternary-Bonsai-2-27B (ternary {−1, 0, +1} weights, 5.95–8.60 GB, claimed 98.2% of FP16 quality) trended #29 on 2026-09-16 via two repos, and Altworld's Hemmingway-1…
Two independent open-weight 27B models built on the Qwen3.8-27B backbone reached #29 on Hugging Face's trending list between 2026-09-16 and 2026-09-20. On 2026-09-16, Prism ML published Ternary-Bonsai-2-27B (27.36B parameters, Apache 2.0) in two repos that both trended #29 minutes apart: the GGUF build (prism-ml/Ternary-Bonsai-2-27B-gguf) and the MLX 2-bit build (prism-ml/Ternary-Bonsai-2-27B-mlx-2bit). The model uses end-to-end ternary {−1, 0, +1} weights in a Hadamard-rotated basis with g128 FP16 scales at 1.72 bits/weight and no FP16 fallback, shrinking the ~54 GB FP16 model to 5.95 GB (PTQ1_0), 7.21 GB (PQ2_0), or 8.60 GB (MLX 2-bit, ~7x reduction). Prism ML claims 98.2% of FP16 quality retained (84.78 average across 14 thinking-mode benchmarks). It keeps a 262K-token context via the base's ~75% linear hybrid attention, ships custom ternary kernels for llama.cpp (CUDA, Metal, CPU) and MLX on Apple Silicon, plus an optional Q8_0 vision tower pack. On 2026-09-20, Altworld released Hemmingway-1, a separate 27B Apache-2.0 Qwen3.8-27B derivative with a 262,144-token context tuned for everyday writing (messages, emails). It claims first place on the vendor's self-built CommunicationBench ahead of Fable 5.1, GPT-6 Astra, Kimi K3, GLM-5.3, Grok 4.6 and DeepSeek V4 Pro, and third on the independent EQ-Bench 4 (within twelve points of the best model); the model card discloses that CommunicationBench, Human-Likeness and StoryBench are the vendor's own benchmarks, run blind and order-balanced with a third-party judge. The reports do not conflict: Report 1's ~5.95 GB figure refers to the PTQ1_0 GGUF packing while Report 2's 8.60 GB refers to the MLX 2-bit pack of the same model.
- Prism ML Ternary-Bonsai-2-27B: 27.36B-parameter model derived from Qwen3.8-27B, Apache 2.0 licensed; trended #29 on 2026-09-16 with two separate repos (-gguf and -mlx-2bit) minutes apart.
- End-to-end ternary {−1, 0, +1} weights in a Hadamard-rotated basis, g128 FP16 scales at 1.72 bits/weight, no FP16 fallback.
- Size reduction from ~54 GB FP16 to 5.95 GB (PTQ1_0 GGUF), 7.21 GB (PQ2_0 GGUF), or 8.60 GB (MLX 2-bit for Apple Silicon) — roughly a 7x reduction.
- Vendor-claimed 98.2% of FP16 quality retained: 84.78 average across 14 thinking-mode benchmarks.
- 262K-token on-device context enabled by Qwen3.8-27B hybrid attention (~75% linear); custom ternary kernels for llama.cpp (CUDA, Metal, CPU) and MLX; optional Q8_0 vision tower pack; whitepaper, demo repo, and forked llama.cpp/MLX runtimes.
- Altworld Hemmingway-1: separate 27B Apache-2.0 model built on Qwen3.8-27B with 262,144-token context, tuned for everyday writing (messages, emails); trended #29 on 2026-09-20.
- Hemmingway-1 claims #1 on vendor-built CommunicationBench, ahead of Fable 5.1, GPT-6 Astra, Kimi K3, GLM-5.3, Grok 4.6 and DeepSeek V4 Pro; placed third on independent EQ-Bench 4, within twelve points of the best model.
- Altworld discloses that CommunicationBench, Human-Likeness and StoryBench are its own benchmarks, claimed to be run blind and order-balanced with a third-party judge.
Coverage timelineoldest first · each row is one article
- · 10d agoprism-ml/Ternary-Bonsai-2-27B-gguf — new model trending #29 on Hugging Face
Hugging Face trending models· 55
Prism ML released Ternary-Bonsai-2-27B, a 27B ternary-weight model derived from Qwen3.8-27B that runs full reasoning in ~5.95 GB GGUF.
- · 10d agoprism-ml/Ternary-Bonsai-2-27B-mlx-2bit — new model trending #29 on Hugging Face
Hugging Face trending models· 48
Prism ML releases Ternary-Bonsai-2-27B, a 27B ternary-weight model derived from Qwen3.8-27B, running on laptops at 8.6 GB with 98.2% FP16 quality retained.