autotrust released an unofficial pruned NVFP4 GLM-5.3-Flash build that fits two NVIDIA DGX Sparks.
autotrust/GLM5.3-Flash-E224-DGX-Spark is an unofficial compact derivative of zai-org/GLM-5.3-Flash for NVIDIA DGX Spark and other Blackwell systems. Neural architecture search keeps 224 of 288 routed experts per layer, experts are stored in NVFP4, and the model still activates 18 billion parameters per token, with weights around 141 GiB. It retains the 154,880-token vocabulary, the vision tower, and the original MTP speculative-decoding layer (about 1.85× single-stream decode). On one B200, the card reports GPQA-Diamond 90.9%, AIME 2025 pass@1 88.3%, HumanEval 98.2% sampled, and MMMU 73.6%, close to unpruned NVFP4 references; Spark memory figures are calculated, not measured on Spark.
Keeps 224 of 288 routed experts and activates 18 billion parameters per token.
NVFP4 weights are 141 GiB, fitting two DGX Sparks with KV-cache headroom.
GPQA-Diamond reaches 90.9% and AIME 2025 pass@1 reaches 88.3% at max effort.
Original MTP layer remains for about 1.85× speculative decode; vision tower is intact.
Accuracy was measured on one B200, not on DGX Spark hardware.
Full article2,446 words · extracted from huggingface.co · click to collapse
# GLM5.3-Flash-E224-DGX-Spark
**autotrust/GLM5.3-Flash-E224-DGX-Spark is a compact build of [zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash) for desktop Blackwell systems like NVIDIA DGX Spark. It's an unofficial derivative.**
It keeps 224 of the 288 routed experts in each layer by **Neural Architecture Search (NAS)**, uses NVFP4 for the experts and still activates 18 B parameters per token. The weights take 141 GiB, which is small enough for **two DGX Sparks connected by ConnectX-7** (256 GB of unified memory in total) with room left for long-context KV cache. A single 180 GB Blackwell GPU (B200/GB200) can also run it.
The original GLM-5.3-Flash MTP layer ships unmodified in [`mtp/`](mtp/) as an optional speculative-decoding draft. It gives about 1.85× single-stream decode speed at the same output quality.
The model keeps the full 154,880-token vocabulary and has the vision tower intact.
| Speculative decoding (MTP) | ✅ | ✅ | ✅ original MTP layer included (`mtp/`, optional) |
## Designed for DGX Spark
DGX Spark (GB10 Grace Blackwell) has 128 GB of LPDDR5X unified memory, 273 GB/s bandwidth and native FP4 tensor cores. That profile shaped this build:
* **Memory budget.** Two Sparks have 256 GB together, but each node has to fit its share of the weights, the KV cache, CUDA graphs, the OS and the desktop. The unpruned NVFP4 checkpoint leaves almost no room for KV cache on each node. At 141 GiB, this model leaves roughly 40 GiB per node. That's enough for 128 K+ thinking traces at `reasoning_effort=max`.
* **Bandwidth budget.** Decode on Spark is memory-bandwidth bound. Every token still reads the same 18 B active parameters (top-8 experts + shared expert + attention), so per-token cost doesn't change. With fewer experts resident, more of the memory stays free for KV cache and batching.
* **FP4 native.** Routed experts use the modelopt NVFP4 format (16-element groups, e4m3 group scale, fp32 tensor scale). GB10's Blackwell tensor cores execute it natively. Attention, shared experts, embeddings and the vision tower are in BF16.
* **Same quality class as the full model.** On every benchmark we measured, the gap to the unpruned model is within noise, except for a few points on GPQA-Diamond (see below).
> ⚠️ **Hardware validation status.** All accuracy and throughput numbers below were measured on a **single NVIDIA B200** with the same weights. The 2× DGX Spark deployment recipe below follows NVIDIA's standard two-Spark vLLM setup. Memory figures for Spark are calculated from the measured weight footprint, not measured on Spark hardware. Expect much lower absolute tokens/s on Spark than on B200, because GB10 has about 30× less memory bandwidth. Reports from Spark owners are very welcome in the Community tab.
## Benchmarks
Measured on one B200 with vLLM. Sampling follows the base model's official recipe (`temperature=1.0, top_p=0.95`); HumanEval also uses greedy decoding. `reasoning_effort` is the GLM-5.3-Flash chat-template thinking budget (`low` / `high` / `max`).
**Scoring is strict:** a response that runs out of tokens before giving a final answer counts as wrong. All numbers are single runs. MoE decoding in vLLM isn't bit-deterministic, so treat ±2–3 points as noise.
Token usage per question (completion tokens, thinking included):
| | mean | median | p90 | max |
|---|---|---|---|---|
| GPQA-Diamond, effort=low | 6.3 K | 0.4 K | 24 K | 65.5 K (budget) |
| GPQA-Diamond, effort=max | 19.7 K | 7.0 K | 53 K | 163.8 K (budget) |
| AIME 2025, effort=high | 15.2 K | 2.4 K | 65.5 K | 65.5 K (budget) |
| AIME 2025, effort=max | 34.2 K | 12.6 K | 145 K | 163.8 K (budget) |
| C-Eval, effort=low | 0.26 K | 0.15 K | 0.3 K | 8.2 K |
**On Spark, size `--max-model-len` for the effort you use.** About 64 K is enough for `low`. Use ≥ 160 K for `max`; the long tail of hard problems runs past 130 K tokens.
BFCL was run with `bfcl-eval` v4 against the local OpenAI-compatible endpoint (`--tool-call-parser glm47`, template-default effort, `--num-threads 32`). `multi_turn_miss_func`, `miss_param`, `long_context` and the agentic web-search/memory categories weren't run.
### Throughput (single B200, reference only)
`reasoning_effort=low`, 1,024 output tokens, short prompts, full CUDA graphs:
These numbers come from a B200 with about 8 TB/s of HBM bandwidth. A DGX Spark has 273 GB/s per node, so single-stream decode there will be much slower. Plan for interactive single-user or small-batch serving on Spark, not high-concurrency throughput.
### MTP speculative decoding (optional)
The `mtp/` folder holds the original GLM-5.3-Flash MTP layer, unchanged: BF16, all 288 experts, 17 GB. vLLM loads it as a separate draft model. Speculative decoding is lossless; accuracy measured with MTP on matches the runs without it, within noise.
Measured on a single B200 with `reasoning_effort=low`, 1,024 output tokens, served from this repository as uploaded. Text, Chinese, tool-calling and image smoke tests pass both with MTP on and off.
**When to use MTP:** turn it on for interactive, low-concurrency serving (1–8 streams), which is the typical DGX Spark workload. Turn it off for high-concurrency batch serving. At 32 concurrent streams, MTP lowers aggregate throughput by about 16 % and raises median TTFT from 0.8 s to 6.4 s, because the draft's extra weights shrink the KV cache: on one B200 at `--gpu-memory-utilization 0.97`, the KV cache drops from 1.31 M to 0.32 M tokens.
With MTP on (2 draft tokens), HumanEval greedy scored 97.0 % and GPQA-Diamond (low) scored 77.8 %, in line with the runs without MTP. Acceptance over the reasoning-heavy eval traffic was 2.36 tokens per step.
**MTP is especially useful on DGX Spark.** Decode there is memory-bandwidth bound, and single-user interactive use is the typical workload, which is exactly where speculative decoding helps most. The cost is 17 GB of extra weights (about 8.5 GB per node at TP=2), so leave room for it in your memory budget (see Deployment).
### Where it loses vs. the unpruned model
* **GPQA-Diamond at low effort:** about 78 % here vs. the low-80s for larger builds. At full budget (`max`), the gap closes to within noise.
* **Vision:** MMMU val 73.6 %. We didn't measure the unpruned model in the same harness, so we can't quantify the gap.
## Deployment
### Requirements
* vLLM with GLM-5.3-Flash support: **vLLM ≥ 0.30.0**, or the official `vllm/vllm-openai:glm53-flash` image. Validated here on the `ZJY0516/vllm@glm-release` branch ([vllm-project/vllm#53906](https://github.com/vllm-project/vllm/pull/53906)) at commit `7e2d791` plus `8f8cc41` ("Make GLM-5.3 kpool metadata graph-safe"). Without that fix, full CUDA graphs can crash under concurrency.
* `transformers >= 5.16.1`
* Set `VLLM_USE_DEEP_GEMM=0`.
### 2× DGX Spark (target configuration)
1. Connect the two Sparks with a QSFP cable on the ConnectX-7 ports. Then follow NVIDIA's ["Connect two Sparks"](https://build.nvidia.com/spark) playbook for networking and passwordless SSH.
2. Download this repository to the same path on **both** nodes.
3. Launch vLLM with tensor parallelism across the two nodes. With the multiprocessing backend, run:
```bash
# on both nodes: point NCCL/Gloo at the ConnectX-7 interface (check `ibdev2netdev`)