HiDream launches HiDream-O1-Video-1.0, an omnimodal video model that debuts at No. 4 on Artificial Analysis image-to-video leaderboard
- HiDream.ai launched HiDream-O1-Video-1.0 (HiDream V1), a native omnimodal video model that takes text, images and video as input and outputs 1080p clips of 5 to 20 seconds with natively synchronized audio.
- The model placed No. 4 globally on the Artificial Analysis Image to Video Leaderboard (With Audio) and No. 8 on the Arena.ai Image-to-Video leaderboard in its first showing on both.
- HiDream V1 runs a three-stage pipeline: global planning of narrative and character states, joint generation that constrains visuals, motion and semantics, then alignment through multimodal reward signals for visual quality, continuity and physical plausibility.
- Before generating, the model parses a prompt across shot duration, setting, character state, movement, facial expression, composition, camera motion, dialogue and ambient sound, a workflow HiDream calls understand, plan and generate.
- CTO Yao Ting said the next generation of video models will not be defined by resolution or duration alone but by whether a model understands creator intent and how objects, actions and sounds interact in the real world.