HiDream launches HiDream-O1-Video-1.0, debuts at No. 4 on Artificial Analysis image-to-video leaderboard with audio
- HiDream.ai released HiDream-O1-Video-1.0 (HiDream V1), a native omnimodal video model that takes text, image, and video inputs and outputs 1080p clips of 5 to 20 seconds with natively synchronized audio.
- On its first two third-party benchmarks, HiDream V1 placed No. 4 on the Artificial Analysis Image to Video Leaderboard (With Audio) and No. 8 on the Arena.ai image-to-video leaderboard, behind Chinese models such as Seedance and MiniMax H3.
- Generation runs through a three-stage framework: narrative and character planning first, then joint constraint of visuals, motion, semantics, and audio, then alignment with multimodal reward signals for continuity and physical plausibility.
- Duration is folded into narrative planning, so the model picks its own length within the 5 to 20 second range based on how an action unfolds, instead of the user setting a fixed window that gets padded with pauses or slow motion.
- The model folds gravity, inertia, collisions, deformation, materials, lighting, and spatial continuity into generation, an attempt to move from frame-level fidelity to modeling how a dynamic world behaves over time.