Two Hugging Face trends: GLM-5.3 quant and Kolibri-1
Hugging Face trends include an EXL3 quant of uncensored GLM-5.3 and Aleph Alpha’s Apache 2.0 Kolibri-1.
The two reports cover separate Hugging Face trending releases, not conflicting accounts of one model. On 1 October 2026, Infatoshi’s EXL3 3.0 bpw quantization of dealignai/GLM-5.3-UNCENSORED-FP8, a no-fine-tune weight edit of Z.ai’s GLM-5.3, was trending at #30. The 753-billion-parameter MoE (256 routed experts, eight active, GlmMoeDsaForCausalLM, MLA attention) is 273 GiB at a 3.04 average bitrate; versus the FP8 source, KL divergence is about 0.09 and perplexity is 3.440 versus 3.302, while tau2-bench airline and retail scores trail stock GLM-5.3 FP8 by 0.070 and 0.022, within sampling error. The card also notes a TabbyAPI parser bug that can turn numeric-looking string tool arguments into integers. On 2 October 2026, Aleph Alpha’s Kolibri-1, an Apache 2.0 German-and-English reasoning model with tool calling, was trending at #28, with 78,103,074,560 total parameters and 3,457,573,120 active per token. It uses 384 experts per layer (one shared, six routed), float8 weights of about 78 GB, a 262,144-token native context checked to 1,048,576, a knowledge cutoff of 18 June 2026, and about 6.4e23 FLOPS after 20 trillion pre-training tokens on 768 NVIDIA B200 GPUs over about 21 days plus 3.44 trillion mid-training tokens.
- On 2026-10-01, Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw was trending #30: an EXL3 quant of dealignai/GLM-5.3-UNCENSORED-FP8, a no-fine-tune weight edit of Z.ai GLM-5.3, at 273 GiB and a 3.04 average bitrate.
- That checkpoint is a 753-billion-parameter MoE (256 routed experts, 8 active) using GlmMoeDsaForCausalLM and MLA attention.
- Versus the FP8 source, WikiText KL divergence is about 0.09 and perplexity is 3.440 versus 3.302; tau2-bench airline and retail trail stock GLM-5.3 FP8 by 0.070 and 0.022, within sampling error.
- A TabbyAPI tool parser bug can coerce numeric-looking string tool arguments to integers.
- On 2026-10-02, Aleph-Alpha/Kolibri-1 was trending #28: Apache 2.0, 78,103,074,560 parameters with 3,457,573,120 active per token, 384 experts per layer (one shared, six routed), float8 weights in 128 by 128 blocks, about 78 GB.
- Kolibri-1 was pre-trained on 20 trillion tokens on 768 NVIDIA B200 GPUs for about 21 days, then 3.44 trillion mid-training tokens and a long-context stage, for about 6.4e23 FLOPS.
- Native context is 262,144 tokens, quality-checked to 1,048,576; knowledge cutoff is 18 June 2026; it supports reasoning and tool calling in German and English.
Coverage timelineoldest first · each row is one article
- · 3d agoInfatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw — new model trending #30 on Hugging Face
Hugging Face trending models· 52
Infatoshi published a 3.0 bpw EXL3 quant of uncensored GLM-5.3, a 753B MoE now trending on Hugging Face.
- · 2d agoAleph-Alpha/Kolibri-1 — new model trending #28 on Hugging Face
Hugging Face trending models· 72
Aleph Alpha released Kolibri-1, an Apache 2.0 78B MoE reasoning model for German and English.