MiniMax Launches H3, an Omni-Modal Video Model Generating 15-Second 2K Clips With Native Stereo Audio
AI video model releases
MiniMax Launches H3, an Omni-Modal Video Model Generating 15-Second 2K Clips With Native Stereo Audio
MiniMax H3 launched July 31, 2026 via API and the Hailuo AI app, generating 2K video from 4 to 15 seconds (integer values only) with native stereo audio, unifying text-to-video, image-to-video, and editing into one pretraining paradigm.
H3-VAE, a rebuilt tokenizer, delivers a stated 4x gain in effective sequence length, cutting training and inference costs enough to make native 2K output economically viable.
In-Context Regeneration replaces bolt-on super-resolution by having the base model re-read its own multimodal context to recover small text and fine details, useful for brand and product rendering.
Third-party trackers report a 2K pay-as-you-go rate of $0.13 per second (about $1.95 for a 15-second clip), though MiniMax's own pricing page still listed only Hailuo 2.3 tiers at time of writing.
Artificial Analysis ranks H3 first in video editing, but behind Google's Gemini Omni Flash in text-to-video and behind both Seedance 2.0 and Gemini Omni Flash in image-to-video; open weights are promised 'in the coming days' but have not shipped.