MiniMax H3’s public vLLM serving stack renders 10.125 seconds of audiovisual video in under nine seconds
- vLLM-Omni benchmarked MiniMax H3 at under nine seconds to produce a completed 10.125-second audiovisual MP4 on eight B300 GPUs, faster than the clip's playback time.
- MiniMax released H3 under a community license in August 2026, and the open-weight audiovisual model now has a public serving path through vLLM-Omni and FastH3.
- FastH3, developed through FastVideo with Hao AI Lab, Nuva Lab, and Nvidia's FastGen team, reduces H3's denoising process from 49 evaluations to four DiT calls.
- The benchmark measures delivery of a finished H.264/AAC MP4 after prompt encoding, joint video-audio latent generation, decoding, transfer, and packaging; it does not measure first-frame latency or prove continuous frame streaming.