HiDream launches HiDream-O1-Video-1.0, an omnimodal video model debuting at No. 4 on Artificial Analysis
- HiDream.ai released HiDream-O1-Video-1.0 (HiDream V1), a native omnimodal video model that takes text, image and video inputs and outputs 1080p clips of 5 to 20 seconds with natively synchronized audio.
- On debut the model ranked No. 4 on the Artificial Analysis Image to Video Leaderboard (With Audio) and No. 8 on the Arena.ai Image-to-Video leaderboard.
- The generation pipeline runs in three stages: global planning of narrative and character states, joint generation that constrains visuals, movement and semantics together, then alignment using multimodal reward signals for visual quality, continuity and physical plausibility.
- Before generating, the model parses a prompt into shot duration, setting, character state, movement, facial expression, composition, camera motion, dialogue and ambient sound, a workflow HiDream calls understand, plan and generate.
- CTO Yao Ting said the next generation of video models will not be defined by higher resolution or longer duration but by whether the model understands creator intent and how objects, actions and sounds interact in the real world.