FastVideo releases open-weight 4-step FastH3 Preview v1 for MiniMax H3 video and audio generation
- FastVideo, Nuva Lab, and NVIDIA FastGen released FastH3 Preview v1, an open-weight 4-step sparse-distilled student of MiniMax H3 for text-to-video-and-audio generation.
- The recommended VSA/Data-Free checkpoint uses four DiT calls, trains from prompts without target videos, and uses 90% sparse Video Sparse Attention with 64-token tiles.
- FastVideo reports up to 14x speedup on one NVIDIA Blackwell GPU, while eight B200 GPUs generate a 15-second 768p video in under 13 seconds.
- The checkpoint supports custom height and width values in multiples of 32, with training and validation across square, portrait, landscape, and ultrawide 768p formats.
- Preview v1 supports T2VA only; first/last-frame conditioning and reference-to-video require separate distilled checkpoints that FastVideo says are under development.