HiDream launches HiDream-O1-Video-1.0, an omnimodal 1080p video model with synced audio
- Beijing-based HiDream.ai released HiDream-O1-Video-1.0 (HiDream V1), a video generation model that takes text, image and video inputs and outputs 1080p clips of 5 to 20 seconds with natively synchronized audio.
- In its first benchmark runs the model placed No. 4 on the Artificial Analysis Image to Video Leaderboard (With Audio) and No. 8 on the Arena.ai image-to-video leaderboard.
- The training framework runs in three stages: global narrative and character-state planning first, then joint generation that constrains visuals, motion and semantics together, then alignment with multimodal reward signals for visual quality, continuity and physical plausibility.
- CTO Yao Ting said the model was built to represent text, video and audio in one unified framework, arguing the next generation of video models will be judged on intent understanding and on how objects, actions and sounds interact rather than on resolution or clip length.
- HiDream positions the release as part of China's continued strength on international video generation leaderboards, a market the announcement describes as increasingly competitive.