HiDream launches HiDream-O1-Video-1.0, an omnimodal video model debuting at No. 4 on Artificial Analysis
- Beijing-based HiDream.ai released HiDream-O1-Video-1.0 (HiDream V1), a native omnimodal video model that accepts text, images and video as input and outputs 1080p clips of 5 to 20 seconds with natively synchronized audio.
- On its first independent evaluations the model placed No. 4 on the Artificial Analysis Image to Video Leaderboard (With Audio) and No. 8 on the Arena.ai image-to-video leaderboard, which HiDream cites as third-party validation of the omnimodal approach.
- The pipeline runs in three stages: global planning of narrative and character states, joint generation that constrains visuals, movement and semantics together, then alignment of quality, continuity and physical plausibility using multimodal reward signals.
- CTO Yao Ting said the goal is understanding creator intent and how objects, actions and sounds interact, rather than higher resolution or longer duration, and the release names Seedance and MiniMax H3 as the Chinese models already holding international leaderboard positions.
- The announcement gives no release date, pricing, access terms or open-weights details for HiDream V1, and the benchmark claims rest on the company's own summary of the two leaderboards.