ZeroHour
Product

MiniMax H3

2 mentions in 7 days · 2 in 30 days · 2 total · first seen · last

Timeline

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Video DeltaNet combines softmax and linear attention for video generation, cutting DiT denoising of a 14.3-second 768p video to 6.7 seconds.

Video DeltaNet (VDN) pairs local softmax attention with a bidirectional linear memory branch using Video Delta Attention, which updates memory once per frame while jointly incorporating its spatial tokens. It is instantiated on MiniMax H3, applying the hybrid to video-to-video interactions while retaining softmax attention for text or audio interactions. With eight-step distillation and an optimized SGLang serving stack, VDN-H3 completes denoising of a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs, a 14.5x speedup over the 50-step dense H3 baseline.

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Video DeltaNet combines local Softmax attention with linear memory, enabling 14.5x faster 768p video generation on eight NVIDIA B200 GPUs.

Video DeltaNet (VDN) is a video-native hybrid attention architecture pairing local Softmax attention with bidirectional linear memory for long-range context, introducing Video Delta Attention that updates memory once per frame with spatial tokens. It is instantiated on MiniMax H3, applying the hybrid to video-to-video interactions while retaining Softmax for text and audio. With eight-step distillation and an optimized SGLang serving stack, VDN-H3 completes DiT denoising for a 14.3-second 768p video in 6.70 seconds on eight NVIDIA B200 GPUs, a 14.5x speedup over the 50-step dense H3 baseline.

Hugging Face daily papersupdated · 22h agofirst · 1d agoAI research 2 sources

Appears with

Entities are extracted by the model from each article. Watching an entity keeps it in this browser only (no account); the watchlist page and dashboard alerts use it.