Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge
Sol-H3 accelerates 33B MiniMax-H3 video diffusion up to 30×, beating real time on an 8×GB200 node.
Sol-H3 is a full-stack inference pipeline that accelerates the 33-billion-parameter MiniMax-H3 video diffusion model on Sol-Engine from cloud NVIDIA GB200 nodes to memory-constrained DGX Spark. A cross-resolution two-stage scheduler builds layout at low resolution and refines detail at high resolution, linked by a learned latent-to-latent map that skips VAE decode-reencode. A recursive self-improvement loop searches kernel fusions and memory layouts for latency and numerical agreement. The stack delivers up to 30× end-to-end speedup and 20% lower memory, generating a 5-second 1344×768 video with audio 3.5× faster than real time on an 8×GB200 node and in under a minute fully resident on one DGX Spark.
- MiniMax-H3 is a 33-billion-parameter open video diffusion model.
- A two-stage scheduler uses low then high resolution without VAE reencoding.
- Recursive self-improvement searches kernel fusions and memory layouts.
- Up to 30× end-to-end speedup and about 20% lower memory.
- A 5-second 1344×768 clip with audio runs 3.5× faster than real time on 8×GB200.
Full article216 words · extracted from arxiv.org · click to collapse
Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substantial computational overhead. Consequently, their practical production is hindered by generation latency in the cloud deployment like NVIDIA-GB200, alongside strict memory limits that pose further challenges at the edge device like DGX-Spark. To address these diverse hardware bottlenecks from cloud to edge device, we present a full-stack inference pipeline that integrates efficient algorithmic design with optimized operator implementations. Algorithmically, we introduce a cross-resolution two-stage generation scheduler that exploits the step-wise nature of diffusion: early low-resolution steps rapidly establish the global layout, while later high-resolution steps focus refinements of local and perceptual details. These stages are connected by a learned latent-to-latent mapping module, completely eliminating the computationally expensive VAE decode-reencode cycle for resolution transferring cross different resolutions. For operator implementation, we deploy a Recursive Self-Improvement (RSI) loop that searches kernel fusions and memory layouts, evaluating latency together with numerical agreement. Together, these optimizations deliver up to 30x end-to-end speedup and 20% lower memory: a 5-second 1344x768 video with audio is generated 3.5x faster than real time on an 8xGB200 node, and in under a minute fully memory-resident on a single DGX Spark.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35110