MiniMax Releases Multimodal Model H3, Generates 15-Second 2K Audio-Video at Under a Third of Rival Costs, Plans Open Source
AI video model releases
MiniMax Releases Multimodal Model H3, Generates 15-Second 2K Audio-Video at Under a Third of Rival Costs, Plans Open Source
MiniMax H3 is a general multimodal model that unifies text, image, video, and audio understanding and generation in one system, generating up to 15 seconds of 2K resolution dual-channel audio-video content.
H3 uses Contextual Omni Representation technology so natural language descriptions (like referencing a video's camera style or applying an audio melody to a person in an image) drive the full generation chain automatically.
With H3-VAE and In-context Regeneration techniques, per-second generation cost at 2K is under one-third of mainstream models, and under half at 768P resolution.
MiniMax will open-source the H3 model weights within days, marking the first time a major domestic Chinese video generation model has released weights publicly, with design compatibility for multiple domestic chips.
MiniMax acknowledges H3 still lags in image fineness and model scale, and plans to merge capabilities with its M series models while scaling further.