Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw — new model trending #30 on Hugging Face
Infatoshi published a 3.0 bpw EXL3 quant of uncensored GLM-5.3, a 753B MoE now trending on Hugging Face.
Infatoshi released an EXL3 3.0 bpw quantization of dealignai/GLM-5.3-UNCENSORED-FP8, a no-fine-tune weight edit of Z.ai's GLM-5.3. The checkpoint is a 753-billion-parameter MoE (256 routed experts, 8 active) using GlmMoeDsaForCausalLM, at 273 GiB and a 3.04 average bitrate. Versus the FP8 source, KL divergence is about 0.09 and perplexity is 3.440 versus 3.302. On tau2-bench, airline and retail scores trail stock GLM-5.3 FP8 by 0.070 and 0.022, within sampling error; a TabbyAPI parser bug can coerce numeric-looking string tool arguments to integers.
- EXL3 3.04 bpw quant of a weight-edited GLM-5.3, 273 GiB.
- 753B MoE: 256 routed experts, 8 active, plus MLA attention.
- WikiText KL about 0.09 versus FP8; perplexity 3.440 vs 3.302.
- tau2-bench airline and retail scores trail stock FP8 within error.
- TabbyAPI tool parser can turn numeric-looking strings into integers.
Full article500 words · extracted from huggingface.co · click to collapse
# GLM-5.3-UNCENSORED EXL3 3.0bpw
EXL3 quantization of [dealignai/GLM-5.3-UNCENSORED-FP8](https://huggingface.co/dealignai/GLM-5.3-UNCENSORED-FP8), itself a weight-edited (no fine-tune) variant of [zai-org/GLM-5.3](https://huggingface.co/zai-org/GLM-5.3). The edit is documented in `CRACK_SURGERY.json` (copied unchanged from the source repo). This repo is not affiliated with dealignai or Z.ai.
- Architecture: `GlmMoeDsaForCausalLM`, 753B total parameters, 256 routed experts (8 active) + 1 shared, MLA attention with DSA sparse indexer, 78 layers + 1 MTP layer
- Average bitrate: 3.04 bpw (`-b 3.0 --hq`; attention and shared experts at 5 bpw, dense MLPs at 4, routed experts at 3), lm_head 6 bpw, `mul1` codebook
- MTP (next-token prediction) layer included (experts 4 bpw, attention and shared expert 6 bpw, uncalibrated), usable as a speculative draft (`draft_mode: mtp` in TabbyAPI)
- Size: 273 GiB
- Converted with ExLlamaV3 at commit `d3739fd`, default calibration (250 rows x 2048 tokens), source read directly from the FP8 checkpoint
## Fidelity vs the FP8 source
`eval/model_diff.py`, 20 rows x 2048 tokens of wikitext-2 test:
| metric | value |
|---|---|
| KL divergence (quant ‖ FP8) | 0.089 |
| KL divergence (FP8 ‖ quant) | 0.097 |
| per-token KL, median / p90 | 0.021 / 0.221 |
| perplexity, quant / FP8 | 3.440 / 3.302 |
| median KL where FP8 top-prob ≥ 0.95 (44% of tokens) | 0.0011 |
## Agentic evaluation
tau2-bench (airline, retail), Pass^1. Agent temperature 1.0, top_p 0.95; user simulator and judges GPT-4.1 at temperature 0. The reference is stock GLM-5.3 (not the uncensored edit) served at FP8 by Z.ai via OpenRouter, run through the same harness. The difference therefore mixes the dealign weight edit and this quantization.
| domain | stock GLM-5.3 FP8 (Z.ai) | this quant | difference |
|---|---|---|---|
| airline (50 tasks x 2 trials) | 0.710 ± 0.045 | 0.640 ± 0.048 | -0.070 (~1.1 SE) |
| retail (114 tasks) | 0.504 ± 0.033 (2 trials) | 0.482 ± 0.047 (1 trial) | -0.022 (~0.4 SE) |
Per trial: airline baseline 0.740 / 0.680, quant 0.600 / 0.680; retail baseline 0.465 / 0.544, quant 0.482. ± is one binomial standard error. Neither difference is statistically significant at these sample sizes; treat the airline point estimate as a possible small regression rather than a measured one.
## Serving
Tested with TabbyAPI on 8x A100 40GB, layer split (`gpu_split_auto`), MTP drafting, 98K-token shared cache.