HiDream launches HiDream-O1-Video-1.0: 1080p, 5-20 second clips with synced audio, debuts at No. 4 on Artificial Analysis
- HiDream.ai released HiDream-O1-Video-1.0 (HiDream V1), a native omnimodal video model that accepts text, image and video inputs and generates 1080p clips of 5 to 20 seconds with natively synchronized audio.
- On its debut it placed No. 4 on the Artificial Analysis Image to Video Leaderboard (With Audio) and No. 8 on the Arena.ai Image-to-Video leaderboard, both third-party evaluations the company cites rather than publishing its own numbers.
- Generation runs in three stages: global planning of narrative and character states, joint constraint of visuals, motion and semantics, then alignment with multimodal reward signals for quality, continuity and physical plausibility.
- The model folds physical factors into generation, including gravity, inertia, collisions, deformation, materials, lighting and spatial continuity, so object movement and environmental response track real-world behavior.
- CTO Yao Ting says the next generation of video models will not be defined by resolution or duration but by whether a model understands creator intent and how objects, actions and sounds interact.