HiDream launches HiDream-O1-Video-1.0, an omnimodal 1080p video model debuting at No. 4 on Artificial Analysis
- HiDream.ai released HiDream-O1-Video-1.0 (HiDream V1), a native omnimodal video model that takes text, image and video inputs and outputs 1080p clips of 5 to 20 seconds with natively synchronized audio.
- The model debuted at No. 4 on the Artificial Analysis Image to Video Leaderboard (With Audio) and No. 8 on the Arena.ai Image-to-Video leaderboard, the only two independent benchmarks cited in the announcement.
- Generation runs through a three-stage framework: global planning of narrative and character state, joint constrained generation of visuals, motion and semantics, then alignment via multimodal reward signals.
- Physical reasoning covers gravity, inertia, collisions, deformation, materials, lighting and spatial continuity, which HiDream positions as the shift from per-frame quality to sequence-level coherence.
- CTO Yao Ting said the model was designed from the outset to represent text, video and audio in one framework, arguing the next generation of video models will not be defined by resolution or duration alone.