Kling AI Releases 3.0 Video and Image Models with 15-Second Clips and Native Multilingual Audio
AI video model releases
Sign in to create alerts.
Kling AI Releases 3.0 Video and Image Models with 15-Second Clips and Native Multilingual Audio
Kling AI launched its 3.0 model series including Video 3.0, Video 3.0 Omni, Image 3.0, and Image 3.0 Omni, available first to Ultra subscribers before a wider rollout.
Video 3.0 extends generation length to 15 seconds and adds native audio generation in English, Chinese, Japanese, Korean, and Spanish, including accent control and multi-character multilingual dialogue scenes.
The models use a unified multimodal architecture combining text-to-video, image-to-video, reference-to-video, and in-video editing into one system, enabling multi-shot storytelling with dynamic camera angle changes.
Video 3.0 Omni builds on the prior Kling Video O1 Elements feature, letting users upload a reference video to extract a character's visual traits and voice for reuse across new scenes, plus a customizable multi-shot storyboard tool specifying duration, shot size, and camera movement.
Image 3.0 and Image 3.0 Omni now output at 2K and 4K resolution, targeting professional use cases like virtual scene visualization and production assets.