HiDream-O1-Video-1.0 debuts at No. 4 on Artificial Analysis image-to-video leaderboard, generating 1080p clips of 5 to 20 seconds with synced audio
- Beijing-based HiDream.ai released HiDream-O1-Video-1.0, a native omnimodal video model that takes text, image and video inputs and generates 1080p clips of 5 to 20 seconds with natively synchronized audio.
- In its debut the model ranked No. 4 on the Artificial Analysis Image to Video Leaderboard (With Audio) and No. 8 on Arena.ai's image-to-video leaderboard, the two independent benchmarks cited in the announcement.
- Generation runs in three stages: global planning of narrative and character state, joint generation that constrains visuals, motion and semantics together, then alignment with multimodal reward signals covering visual quality, continuity and physical plausibility.
- The model folds physical reasoning into generation, including gravity, inertia, collisions, deformation, materials, lighting and spatial continuity, so object movement and environmental responses follow real-world behavior.
- Instead of requiring a fixed duration up front, the model plans content-adaptive duration inside its narrative planning and picks a length within 5 to 20 seconds from how an event unfolds, avoiding trailing pauses or slow motion used to fill time.