autotrust/GLM5.3-Flash-E224-DGX-Spark — new model trending #29 on Hugging Face
autotrust released an unofficial pruned NVFP4 GLM-5.3-Flash build that fits two NVIDIA DGX Sparks.
autotrust/GLM5.3-Flash-E224-DGX-Spark is an unofficial compact derivative of zai-org/GLM-5.3-Flash for NVIDIA DGX Spark and other Blackwell systems. Neural architecture search keeps 224 of 288 routed experts per layer, experts are stored in NVFP4, and the model still activates 18 billion parameters per token, with weights around 141 GiB. It retains the 154,880-token vocabulary, the vision tower, and the original MTP speculative-decoding layer (about 1.85× single-stream decode). On one B200, the card reports GPQA-Diamond 90.9%, AIME 2025 pass@1 88.3%, HumanEval 98.2% sampled, and MMMU 73.6%, close to unpruned NVFP4 references; Spark memory figures are calculated, not measured on Spark.