MiniMax Launches H3, a Multimodal Video Model That Edits Footage and Generates Native Audio
AI video model releases
MiniMax Launches H3, a Multimodal Video Model That Edits Footage and Generates Native Audio
MiniMax released H3, a video model that accepts four input types at once, text, images, video, and audio, reasoning across them together instead of processing each separately.
H3 generates native stereo audio in the same pass as the video, including dialogue, ambient sound, footsteps, and music, with lip sync accurate enough to hold up on fast rap delivery.
H3 can edit existing footage, replacing or removing objects and people, changing backgrounds and lighting, adjusting visual effects, modifying performance, and swapping dialogue or vocal timbre while keeping the rest of the scene unchanged.
H3 currently leads the Artificial Analysis video editing leaderboard, ranking ahead of Seedance 2.0.
The article frames H3's editing capability as the feature that separates it from competing video models, since most rivals regenerate an entire scene rather than making a localized change.