Alibaba puts Wan3.0 video model into beta with 30-second multimodal generation · cho.sh
Alibaba puts Wan3.0 video model into beta with 30-second multimodal generation
AI video model releases
Alibaba puts Wan3.0 video model into beta with 30-second multimodal generation
Alibaba has released Wan3.0 in beta, generating videos up to 30 seconds from text, images, PDFs, web pages, and PowerPoint files, twice Wan2.5's stated maximum length.
The unified multimodal model accepts text, images, video, and audio in one prompt, with limits of 10 images, five videos, and five audio clips.
Alibaba says Wan3.0 reduces visual drift by preserving reference character identities, props, and spatial layouts across a clip, though independent benchmarks are not yet available.
Wan3.0 is available on wan.video, Alibaba Cloud Model Studio, and the Qwen Cloud API; a 30-second 1080p video costs $6.00 on Standard or $8.40 on Prime.
The beta includes automatic length recommendations and a clip-extension tool, while Alibaba positions document-to-video generation for marketing, training, film production, and simulation footage.