MiniMax H3 unifies text, image, video and audio into one AI model with native 2K output
AI video model releases
MiniMax H3 unifies text, image, video and audio into one AI model with native 2K output
MiniMax H3 launched July 31, 2026, merging text, image, video and audio generation into a single model that outputs 2K clips of 4 to 15 seconds with native stereo sound.
The model replaces separate specialist tools (text to video, image to video, first/last frame, subject reference, motion reference, video editing) with one system that reads all inputs in a shared context window.
A rebuilt tokenizer called H3-VAE compresses video data to make 2K output affordable, and an in-context regeneration step redrafts low resolution clips by re-reading the original prompt instead of using a generic upscaler.
MiniMax claims H3's per second cost at 2K is under a third of mainstream rivals; third party trackers estimate pay as you go pricing near $0.13 per second, or about $1.95 for a 15 second 2K clip, though MiniMax has not confirmed this on its own pricing page.
Independent benchmarks rank H3 first for video editing and top three for text to video and image to video, but it still trails Google's Gemini and ByteDance's Seedance 2.0 in several categories, and an open weights release is promised but not yet shipped.