MiniMax releases omni-modal H3 video model with 2K, 15-second clips and native stereo audio, plans open weights
MiniMax releases omni-modal H3 video model with 2K, 15-second clips and native stereo audio, plans open weights
AI video model releases
Sign in to create alerts.
MiniMax H3, released July 31, 2026, generates up to 15 seconds of 2K video with native stereo audio from a unified text, image, video and audio context, and is live in the MiniMax API (model ID MiniMax-H3) and the Hailuo AI app.
H3 merges tasks that were separate expert systems in Hailuo 01/02 (text-to-video, image-to-video, subject reference, motion reference, editing) into one pre-training paradigm, letting users describe reference and editing relationships in natural language.
H3-VAE, a rebuilt tokenizer, gives a 4x gain in effective sequence length and enables native 2K output, while H3-In-Context Regeneration replaces super-resolution by having the base model regenerate its own low-res result using the original multimodal input.
MiniMax prices 2K generation at under a third of the per-second cost of mainstream video models, and 768p at under half the cost of mainstream 720p models.
MiniMax says it plans to open the model weights within days (pending legal review), but no Hugging Face repo or model card had appeared at time of writing.