Alibaba puts Wan3.0 video model into beta with 30-second multimodal generation
- Alibaba has released Wan3.0 in beta, generating videos up to 30 seconds from text, images, PDFs, web pages, and PowerPoint files, twice Wan2.5's stated maximum length.
- The unified multimodal model accepts text, images, video, and audio in one prompt, with limits of 10 images, five videos, and five audio clips.
- Alibaba says Wan3.0 reduces visual drift by preserving reference character identities, props, and spatial layouts across a clip, though independent benchmarks are not yet available.
- Wan3.0 is available on wan.video, Alibaba Cloud Model Studio, and the Qwen Cloud API; a 30-second 1080p video costs $6.00 on Standard or $8.40 on Prime.
- The beta includes automatic length recommendations and a clip-extension tool, while Alibaba positions document-to-video generation for marketing, training, film production, and simulation footage.