HiDream launches HiDream-O1-Video-1.0, a native omnimodal video model debuting at No. 4 on Artificial Analysis
- HiDream.ai released HiDream-O1-Video-1.0 (HiDream V1), a native omnimodal video model that accepts text, image and video inputs and outputs high-fidelity 1080p clips of 5 to 20 seconds with natively synchronized audio.
- On its first independent benchmarks the model placed No. 4 on the Artificial Analysis Image to Video Leaderboard (With Audio) and No. 8 on the Arena.ai Image-to-Video leaderboard.
- Generation runs in three stages: global planning of narrative and character states, joint generation that constrains visuals, movement and semantics together, then alignment of visual quality, continuity and physical plausibility using multimodal reward signals.
- CTO Yao Ting said the next generation of video models will not be defined by higher resolution or longer duration, but by whether a model understands creator intent and how objects, actions and sounds interact in the real world.
- The article notes that Chinese teams such as Seedance and MiniMax H3 already hold strong leaderboard positions, and HiDream V1's debut widens China's presence among leading video generation models.