OpenVDN/vdn-minimax-h3 — new model trending #12 on Hugging Face
OpenVDN releases VDN-H3, an open hybrid-attention video model on MiniMax H3 that renders a 14.4-second 768p clip in 11.23 seconds on 8 B200 GPUs.
VDN-Minimax-H3 (VDN-H3) adds a frame-wise linear attention branch plus two LoRA adapters to MiniMax H3, distilled into 8-step and 50-step variants. It generates 768p, 14.4-second clips in 11.23 seconds on 8 B200 GPUs (90.5 seconds on one H200) using 8 denoising steps. Weights (about 82 GB total, including the 72 GB H3 base), the optimized inference stack, and training code are fully open-source under the MiniMax H3 Community License, which excludes the EU, UK, Korea, and US.
- Plug-and-play: linear branch and LoRA adapters merge at inference without touching backbone weights
- 8-step variant distilled from a 50-step model; distributed FP8 inference on B200s and H200s
- Prompts encoded with Qwen3-VL-32B; H3-Context-IR rewriting recommended for quality
- Inference path built on Diffusers, FlashAttention, and Triton
Full article844 words · extracted from huggingface.co · click to collapse
# Video DeltaNet: Hybrid Attention to Speed Up Video Models with Near-Lossless Quality
[[`Blog`](https://openvdn.github.io/)] [[`Code`](https://github.com/OpenVDN/vdn-minimax-h3)] [[`🤗 Weights`](https://huggingface.co/OpenVDN/vdn-minimax-h3)] [[`License`](#license)]
We release **VDN-Minimax-H3** (**VDN-H3**), a hybrid-attention model that
generates video faster than it plays, powered by
[MiniMax H3](https://huggingface.co/MiniMaxAI/MiniMax-H3). It offers these key
features:
- **Fast inference:** On 8 B200 GPUs, VDN-H3 generates a 14.4-second clip in
**11.23 seconds** using 8 denoising steps.
- **Hybrid Architecture:** We propose a hybrid-attention architecture: one frame-wise
linear attention branch that is highly efficient, and a softmax branch that maintains
the backbone's visual quality and consistency.
- **Plug-and-Play:** The checkpoint adds a separate linear attention branch and two
small LoRA adapters that can be merged into the backbone during inference without
touching the backbone weights.
- **Fully open-source:** We don't just open-source the weights. The optimized inference
stack and its corresponding training code are released together.
This repository holds the weights. Read the [Blog](https://openvdn.github.io/) for the
architecture, training method, benchmarks, and qualitative results. The inference and
training code, together with the setup instructions, is available at
[OpenVDN/vdn-minimax-h3](https://github.com/OpenVDN/vdn-minimax-h3).
## News
- **September 6, 2026:** We released the [VDN-H3 blog](https://openvdn.github.io/),
[training and inference code](https://github.com/OpenVDN/vdn-minimax-h3), and
[model weights](https://huggingface.co/OpenVDN/vdn-minimax-h3).
## Download the weights
Download everything (about 82 GB) into `ckpts/` using
```bash
hf download OpenVDN/vdn-minimax-h3 --local-dir ckpts
```
The layout will look like
```text
ckpts/
h3-base/ the released MiniMax-H3: transformer, video and audio VAEs, schedulers · 72 GB
stage-b-step-2000/ VDN-H3-50-step: linear_branch/ + adapters/default/ LoRA · 4.3 GB
stage-dmd-step-250/ VDN-H3-8-step: the above + adapters/turbo/ · 5.1 GB
```
`stage-dmd-step-250` is the 8-step model the headline numbers use;
`stage-b-step-2000` is the 50-step model it is distilled from. Each directory
describes itself: `model_spec.json` records the hybrid architecture, `metadata.json`
the training recipe, and the learned tensors sit in `linear_branch/model.safetensors`
and `adapters/<name>/adapter_model.safetensors`.
## Your first render
First clone and set up the
[code repository](https://github.com/OpenVDN/vdn-minimax-h3#set-up-environment).
Then, from the root of that repository, run the model on a single GPU:
```bash
bash scripts/inference/8nfe_tuned_fp8.sh
```
Note that the first run needs to compile all of the kernels, which might take several
minutes. Later runs can reuse the cache.
For your own prompt, you should first encode it using the Qwen3-VL-32B VLM, then render
it through the main diffusion model:
```bash
python src/inference/encode_prompt.py --prompt "..." --out prompts/mine.pt
python src/inference/infer.py \
--config configs/inference/8nfe_tuned_fp8.yaml \
checkpoint=ckpts/stage-dmd-step-250 \
render.prompt_file=prompts/mine.pt \
render.out=results/mine.mp4
```
We strongly recommend rewriting it first using
[H3-Context-IR](https://platform.minimax.io/docs/api-reference/video-generation-v2-h3-context-ir)
or the official
[prompt-writing skills](https://github.com/MiniMax-AI/MiniMax-H3/tree/main/skills)
before encoding it. This can greatly improve the generated video quality.
## Results
We report steady-state denoising speed on the 768p, 14.4-second video generation
workload for the released model using our inference pipeline on H200s and B200s:
**H200:**
| Configuration | GPUs | Seconds/NFE | 50 NFE (VDN-H3-50-step) | 8 NFE (VDN-H3-8-step) |
|---|---:|---:|---:|---:|
| dense MiniMax-H3 | 1 | 32.7 | 27.3 min | 4.4 min |
| VDN-H3 FP8 | 1 | 11.2 | 9.4 min | 90.5 s |
| VDN-H3 FP8 Distributed | 8 | 2.29 | 1.9 min | 18.3 s |
**B200:**
| Configuration | GPUs | Seconds/NFE | 50 NFE (VDN-H3-50-step) | 8 NFE (VDN-H3-8-step) |
|---|---:|---:|---:|---:|
| dense MiniMax-H3 (cuDNN) | 1 | 16.74 | 13.95 min | 2.23 min |
| VDN-H3 FP8 | 1 | 6.41 | 5.3 min | 51 s |
| VDN-H3 FP8 Distributed | 8 | 1.40 | 1.2 min | 11.23 s |
We exclude model loading, warm-up, VAE decoding, and MP4 encoding. For a live setup,
we recommend running the text prompt rewriter, VAE decoding, and MP4 conversion on
separate machines, so the eight GPUs only perform denoising.
## Acknowledgement
VDN-H3 is built on [MiniMax H3](https://huggingface.co/MiniMaxAI/MiniMax-H3) and starts
from its released transformer weights. We also thank
[Diffusers](https://github.com/huggingface/diffusers),
[FlashAttention](https://github.com/Dao-AILab/flash-attention), and
[Triton](https://github.com/triton-lang/triton), on which the optimized inference path
is built. We thank [Kernel Design Agents (KDA)](https://github.com/mit-han-lab/kernel-design-agents)
for kernel design support. We also thank
[Flash Linear Attention (FLA)](https://github.com/fla-org/flash-linear-attention) and
[FlexAttention](https://pytorch.org/docs/stable/nn.attention.flex_attention.html) for
their open-source attention implementations.
## BibTeX
```bibtex
@misc{xi2026videodeltanet,
title = {VideoDeltaNet on MiniMax H3},
author = {Haocheng Xi and Yiming Xie and Hexu Zhao and Yiwen Zhang and Michael Liu and Thomas Creavin and Kurt Keutzer and Xiuyu Li and Zhaoyang Lv and Chenfeng Xu and Haiwen Feng},
year = {2026},
url = {https://openvdn.github.io/}
}
```
## License
VDN-H3 is a derivative of MiniMax-H3 and is distributed under the
[MiniMax H3 Community License Agreement](LICENSE), included here verbatim from the
upstream repository.
The agreement grants rights only in its applicable territory, defined as worldwide
**excluding the European Union, the United Kingdom, the Republic of Korea, and the
United States of America**. It states that use outside the applicable territory is not
authorized and invites people in an excluded territory to contact MiniMax about
obtaining a license.
The agreement also contains redistribution requirements and an Acceptable Use Policy.
Among other requirements, a distribution must include the agreement, modified files
must carry notices of their modification, and distributions to third parties other than
through hosted services must include the `NOTICE` file supplied with the code
repository. Please read the agreement in full before using or distributing VDN-H3. This
note is not a substitute for the license text or for legal advice.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/OpenVDN/vdn-minimax-h3