MiniMax H3 Launches with Open Weights, Native Stereo Audio, and Day-Zero ComfyUI Support
- MiniMax H3 is MiniMax's third generation video model after Hailuo 01 and Hailuo 02, and the first one released with open weights.
- The model generates video up to 2K resolution and 15 seconds long, with native stereo audio produced in the same pass as the video rather than added afterward.
- H3 supports five workflows in one model: text-to-video, image-to-video, first-and-last-frame control, and reference-to-video for transferring subjects, motion, or voice from supplied images, video, or audio.
- ComfyUI added day-zero native support, and MiniMax says pruning the model's modulation weights (about 40% of parameters) into a lookup table cuts memory footprint 66%, from 123.6GB to 42.5GB, letting it run locally on an RTX 3060.
- HN users report real generation times of about 10 minutes for a 10 second 480p clip on a 16GB 3060 or 4070ti, and 3 minutes on a 5080, with mixed opinions on output quality ranging from "spectacular" to "bland and generic".
Hacker News 의견들
I saw the samples people posted and immediately deleted my LTX2 and WAN folders, those are worthless now compared to this.
On the license side, you basically just pinkie promise you won't upset Disney and they hand you a license, that's the whole EU/UK/US compliance story here.
This is AGI, Hollywood should be on red alert.
The example videos look exactly like other people's highly produced work, I can't believe this wasn't trained on stolen content and there's zero protection for artists.
Normal people already dislike AI content, I see this being used for ads, mockups, propaganda, and pre-viz, not for actual films people pay to watch.
Assuming the memory footprint claim is true, a 15 second clip on a 16GB 3060 takes around 10 minutes.
The mouse render is genuinely impressive, a big leap over current SOTA, and the fact that it's open weights is a massive win for the community.
This is still about a year and a half behind Seedance 2.0/2.5, but it puts price pressure on closed foundation models and lets creators avoid heavy-handed safety filters.
Right before that hiking shot, the breath vapor clouds don't sync with the person's actual breathing, devs need to fix that.
Running this on my 4070ti super with 16GB VRAM, a 10 second 480p clip takes 10 minutes and the results are spectacular.
On a 5080 16GB it only takes 3 minutes for the same 10 second 480p mouse workflow.
I tried a 10 second 1080p clip on a bigger machine and the results were actually pretty poor, unusable.
Impressive tech, but aesthetically it all looks painfully bland and generic to me.
This is the worst the model will ever be starting today, so chill out about the aesthetics.
Not sure if bland output is a model limitation or a prompting issue, I've seen people use insanely detailed prompts to get consistent, interesting results out of similar models.
The claim that modulation weights (about 40% of params) can be pruned into a lookup table with zero quality loss seems almost too simple, wondering if this trick applies to LLMs too.
Modulation weights here means adaLN parameters for layer norm conditioning, general LLMs don't really have those so this trick probably doesn't transfer directly.
This lookup table trick is well known in diffusion models specifically because timestep is bounded 0 to 1, so you can discretize and index modulation params per timestep, it's lossless and diffusion-specific.
Human directors still matter here, they'll just prompt and arrange AI-generated shots the way an EDM producer arranges sounds instead of playing instruments live.
Seedance 2.5 just dropped and is significantly better than this for a lot of cases, but this is the best free option right now.