HiDream launches HiDream V1 omnimodal video model, debuts at No. 4 on Artificial Analysis image-to-video leaderboard
- Beijing-based HiDream.ai launched HiDream-O1-Video-1.0 (HiDream V1), a native omnimodal video generation model that takes text, image and video inputs and outputs 1080p clips of 5 to 20 seconds with natively synchronized audio.
- The model debuted at No. 4 on the Artificial Analysis Image to Video Leaderboard (With Audio) and No. 8 on the Arena.ai image-to-video leaderboard, both third-party rankings the company cites as validation of its approach.
- HiDream V1 uses a three-stage framework: global narrative and character-state planning, then joint constraint of visuals, motion and semantics during generation, then alignment of visual quality, continuity and physical plausibility through multimodal reward signals.
- The model folds physical reasoning into generation, including gravity, inertia, collisions, deformation, materials, lighting and spatial continuity, so object and character motion follows real-world behavior rather than frame-level appearance alone.
- CTO Yao Ting framed the design goal as moving video generation past resolution and duration toward understanding creator intent and how objects, actions and sounds interact, and the company claims content-adaptive duration instead of a fixed preset length.