MiniMax open-sources H3, an omni-modal video model with native stereo audio up to 2K and 15 seconds
MiniMax open-sources H3, an omni-modal video model with native stereo audio up to 2K and 15 seconds
AI video model releases
Sign in to create alerts.
MiniMax H3 is now open source as MiniMax's next-generation general-purpose video model, handling text, images, video, and audio as unified multimodal input.
H3 generates video with native stereo audio at resolutions up to 2K and durations up to 15 seconds, with output audio at 32 kHz stereo and 24 FPS frame rate.
The system ships in two variants: H3-Base-FL2VA for first-and-last-frame video generation (zero, one, or two image inputs), and H3-Base-Ref2VA for omni-reference generation supporting up to 9 images, 3 video clips, and 3 audio clips as combined reference inputs.
H3 supports a wide range of output aspect ratios (21:9, 16:9, 4:3, 1:1, 3:4, 9:16) and stable dialogue generation in 11 languages including English, Chinese, Japanese, Korean, and Arabic.
A dedicated module called H3-Context-IR converts complex multimodal instructions into an intermediate representation before generation, which MiniMax says is critical to output quality.