OpenVDN Releases VDN-H3 Video Patch Reporting 14.4-Second 768p Clips in 11.23 Seconds
- VDN-H3 reportedly generates a 14.4-second 768p video in 11.23 seconds on 8 Nvidia B200 GPUs using 8 denoising steps, which the team says is 75 to 90x faster than dense MiniMax H3.
- The model replaces full video attention with sliding-window softmax for nearby frames and a Video Delta Attention linear branch for long-range context, aiming to retain local detail and temporal consistency.
- OpenVDN releases the Hugging Face checkpoint, training code, and optimized inference code; the 8-step checkpoint uses DMD2 distillation from a 50-step model.
- VDN-H3 is distributed as a linear-attention branch plus two LoRA adapters that merge into a frozen MiniMax H3 backbone at inference, and a community contributor has published a native ComfyUI port.
- The license excludes use in the US, EU, UK, and South Korea, and the reference setup requires Hopper or datacenter Blackwell GPUs.