Alibaba opens Wan3.0 beta with 30-second video generation and document-to-video input

Alibaba opens Wan3.0 beta with 30-second video generation and document-to-video input

  • Alibaba opened Wan3.0 in beta on August 24, raising a single video generation run to 30 seconds and listing long-clip consistency for characters, props, sound, spatial relationships, and visual style as a release capability.
  • Wan3.0 accepts DOC, XLS, PPT, PDF, and Markdown files as inputs, allowing existing office materials to be converted into video without first rewriting them as prompts or shot lists.
  • The API costs 0.3, 0.6, and 1.2 yuan per second at 480P, 720P, and 1080P; a 30-second 1080P generation is 36 yuan, or 25.2 yuan during the August 24 to September 23 30% promotion.
  • Alibaba made Wan3.0 available through Alibaba Cloud Bailian, Wan Jing Yi Ke, the Wanxiang website, and Qwen Chat's desktop client, while rolling it out gradually in the Qwen app.
  • The article says ByteDance Seedance 2.5 had already reached 30-second single clips, while independent benchmarks for Wan3.0 have not yet appeared.
AI video model releases
HiDream launches HiDream-O1-Video-1.0, an omnimodal 1080p video model debuting at No. 4 on Artificial AnalysisFal's H3 Max claims 35x throughput of MiniMax's H3 endpoint, tops two image-to-video leaderboardsAlibaba's Qwen3.8-Omni-Flash cuts video input costs about 89% versus Qwen3.5-Omni-PlusAlibaba releases Qwen3.8-Omni-Flash with 1M-token context, claiming Gemini 3.8 Flash-level audio and videoHiDream launches HiDream-O1-Video-1.0, an omnimodal video model debuting at No. 4 on Artificial AnalysisAlibaba Cloud's Wan3.0 video model outputs a single 30-second long take and accepts up to five video referencesShengShu Technology Unveils Vidu S2, Splitting Real-Time AI Video Into an Interactive Avatar Model and a Live Stream EditorHiDream launches HiDream-O1-Video-1.0, a native omnimodal video model debuting at No. 4 on Artificial AnalysisKling 3.0 launches with native 4K 60fps video, multi-shot storyboarding, and 5-language lip syncShengShu unveils Vidu S2, adding real-time avatar interaction and live stream editingShengshu Technology Launches Vidu S2 With Real-Time Editing, Dynamic Reference Images and VR Headset StreamingByteDance Open-Sources Bernini, an Apache-2.0 Video Editor That Plans the Edit Before It RendersByteDance Restricts Seedance 2.0 Video Model Three Days After Release, Following Viral Deepfake DemoKuaishou launches Kling 3.0 with native 4K 60fps, 6-shot AI Director, and 5-language lip syncJD.com open-sources EchoWM interactive audio-visual world model and shows JoyAI-Echo1.5 clearing 10 minutes of continuous generationKuaishou's Kling AI launches Kling 3.0 with native 4K 60fps video, multi-shot storyboarding, motion control and 5-language lip syncBlack Forest Labs releases FLUX Video Edit, an API-only video editing model at $0.03 per secondKling AI Ships Kling 3.0: Native 4K 60fps, Multi-Shot Storyboarding, 5-Language Lip SyncKuaishou Says Kling 3.0 Generates 4K 60fps Video With Six-Shot SequencesOpenVDN Releases VDN-H3 Video Patch Reporting 14.4-Second 768p Clips in 11.23 SecondsKuaishou Announces Kling 3.0 With Claimed Native 4K 60fps and Six-Shot Video GenerationTwelve Labs' Marengo becomes Amazon Bedrock's first video modelJD.com Open-Sources JoyAI-EchoWM for Interactive Video and Audio World Generationfal releases H3 Max, claiming top image-to-video rankings and 3-second generation for 5-second clipsOverworld releases Waypoint-1.5 real-time world model with 720p 60 FPS and 360p tiersAdobe puts Firefly, Veo, Runway, Kling and Luma video generation inside Premiere timelinesNVIDIA Releases Sol-H3 Inference Code, Running MiniMax H3 Video Faster Than Playback on Eight B300 GPUsNVIDIA and SANA Release Sol-H3 Stack That Generates 5-Second MiniMax-H3 Video in 1.653 SecondsCreateAI open-sources Ruyi-Mini-7B image-to-video model for consumer GPUsWorld Labs says Atlas generates camera-controlled views and 3D scenes from sparse imagesLuma releases Ray3 video model with self-evaluation claims and 16-bit HDR EXR outputCaira camera app adds Gemini Omni and Runway Aleph video editing in private betaGoogle launches agentic video understanding for Gemini Flash, claiming up to 88% fewer video tokensWorld Labs unveils Atlas, a 1440p video model with pixel-level camera controlSand.ai open-sources 114B-parameter MAGI Preview video model with 6B active MoE parametersWorld Labs unveils Atlas, a spatial world model that generates 1440p videos up to one minutexAI's Grok Imagine Raises Reference Limit From 7 to 14 for Video GenerationRunway previews GWM Worlds 2 for real-time interactive 720p video and audio worldsfal releases H3 Max, a post-trained MiniMax H3 video model claiming #1 benchmark rankingsGoogle adds agent-based video retrieval to Gemini Flash, claiming 88% lower token usefal releases H3 Max, a post-trained MiniMax H3 video model claiming 3-second generation for 5-second clipsWorld Labs releases Atlas for controllable 1440p, one-minute 3D world generationfal Releases H3 Max, a Post-Trained MiniMax H3 Video Model Claiming #1 Benchmark RanksGoogle Adds Agentic Video Understanding to Gemini Flash Modelsfal Releases H3 Max, a Post-Trained MiniMax H3 Video Model That Generates Five Seconds in About ThreeWorld Labs releases Atlas world model with 3D reconstruction and camera trajectory controlfal releases H3 Max, a post-trained MiniMax H3 video model that generates 5-second clips in about 3 secondsGoogle adds agentic video analysis to Gemini Flash models, claiming 88% lower token useMiniMax H3’s public vLLM serving stack renders 10.125 seconds of audiovisual video in under nine secondsWorld Labs launches Atlas, a shared model for controlled video, 3D reconstruction and robot viewsVisko launches Orbis for continuous real-time video generation with live prompt changesRunway Solaris Generates Stateful App Interfaces Frame by Frame Without CodeWorld Labs unveils Atlas, a multimodal world model that generates controllable 1440p videos up to one minuteGoogle adds agentic video understanding to Gemini Flash models, claiming up to 88% lower token useVidu Announces S1 Voice-Driven Real-Time Interactive Video ModelVisko opens Orbis, a real-time 4K interactive video model, after $10M pre-seed roundVisko opens Orbis, a live AI video model that streams interactive 4K worlds for hoursfal releases H3 Max, claiming top video benchmark rankings and 3-second generation for 5-second clipsFastH3 open-weight preview cuts MiniMax H3 video and audio generation to four DiT callsGemini Omni 1.1 Flash Adds 40-Second Scene Extension and First-to-Last-Frame Video ControlAlibaba Wan 3 unifies four video modes, adds 30-second generation and native audioOverworld releases Waypoint-1, a real-time video diffusion model controlled by text, mouse, and keyboardNVIDIA releases Cosmos 3 open multimodal world model for physical AIAlibaba launches Wan3.0 AI video model for 30-second document-based clipsRunway releases Gen-4.5, claiming top rank on Artificial Analysis' text-to-video benchmarkGoogle releases Veo 3.1 Lite in paid preview at under half the cost of Veo 3.1 FastFastVideo releases open-weight 4-step FastH3 Preview v1 for MiniMax H3 video and audio generationReported Gemini Omni 1.1 Flash update chains 10-second clips into 40-second scenes with 4K outputGoogle's Gemini Omni 1.1 Flash adds 40-second video extension, frame pinning, and 4K outputGoogle releases Gemini Omni 1.1 Flash with 40-second scene extension and 4K upscalingGoogle Gemini Omni 1.1 Flash adds 4K output, 40-second scene extension, and keyframe controlsGoogle releases Gemini Omni 1.1 Flash with video continuation, frame controls, and 4K upscalingGoogle makes Gemini Omni 1.1 Flash generally available with scene extension and 4K outputAlibaba releases Wan3.0, a 30-second video model that takes documents and web pages as inputsZ.ai open-sources 320B-parameter GLM-5.3-Flash for million-token video inputsGoogle Omni 1.1 Flash adds 1080p and upscaled 4K video outputGoogle's Gemini Omni 1.1 Flash Reportedly Adds Direct 4K Video Output and 40-Second ExtensionsGemini Omni 1.1 Flash adds 4K upscaling, 40-second scene extension, and frame-to-frame video generationFastVideo releases FastMetal-QAD video models for local Apple Silicon generationGoogle releases Gemini Omni 1.1 Flash with 40-second scene extension and 4K video upscalingAlibaba opens Wan3.0 beta with 30-second video generation and document-to-video inputGoogle releases Gemini Omni 1.1 Flash with 40-second scene extension and 4K upscalingJD.com Open Sources Echo-WM, a Navigable Video, Audio, Music and Speech World ModelAlibaba WAN 3.0 raises single-pass video length to 30 seconds, adds document and web inputsAlibaba releases Wan3.0 video model after $10 billion share placementAlibaba Releases Closed-Weight Wan3.0 Video Model With 30-Second 1080p GenerationAlibaba's Wan3.0 adds 30-second video, audio, and document inputs at $6 per 1080p clip4DAnyone generates 16 consistent views from one video for 4D human reconstructionAlibaba's Wan 3.0 beta claims 30-second single-pass AI video and document inputsAlibaba rolls out Wan 3.0 with 30-second 1080p document-to-video generationAlibaba releases Wan3.0 with 30-second generation and document-to-video inputAlibaba Rolls Out Wan3.0 Video Model With 30-Second Generation and Document InputsAlibaba launches Wan3.0 video model with 30-second generation and document inputsLightricks Releases Apache 2.0 LTX-2.3 Open-Weight 4K Video ModelAlibaba launches closed Wan3.0 video model with 30-second 1080p generationAlaya Lab releases Evoke, a 14B open world model with external scene memory for hour-long interactionAlibaba puts Wan3.0 video model into beta with 30-second multimodal generationAlibaba launches Wan3.0 with 30-second video generation from documents and slide decksLTX releases open-weight LTX-2.5 video world model with multishot generation and 6.8-second 720p clipsAlibaba launches Wan3.0, which turns business documents into 30-second videos