Alibaba’s ‘Seedance’ rival Wan3.0 intensifies AI video race

  • New model adds document-to-video generation and longer clips for real-world content creation
  • Alibaba joins global AI video race as models move from short clips to production tools

Alibaba Cloud has opened public testing for Wan3.0, the latest version of its AI video generation model, as the company pushes the technology beyond short visual clips toward longer, more practical content creation.

Compared with its predecessor Wan2.7, the new model upgrades video generation across three areas: duration, input formats and visual realism.

Alibaba said Wan3.0, available for public beta from August 6, can generate videos of up to 30 seconds in a single request, marking a shift from generating isolated shots to creating more complete visual narratives.

The company said only two AI video models globally currently support single-generation videos of up to 30 seconds.

From prompts to production materials

A major upgrade in Wan3.0 is its expanded input capability.

Beyond text, images, audio and video, the model now supports common document formats including DOC, XLS, PPT, PDF and Markdown files. Users can directly convert teaching materials, product presentations and business reports into videos.

Each uploaded file can be up to 100MB and contain up to 50 pages, allowing business and educational content to become video outputs without requiring users to manually write prompts or rebuild materials.

More realistic characters and scenes

Alibaba said Wan3.0 has also improved its ability to generate realistic human characters and complex scenes.

The model produces more detailed facial features and skin textures, while group scenes can display different characters with varied emotions and expressions.

Its reference-based generation capability allows users to maintain consistency across key elements including characters, objects, environments and visual styles — a major challenge for current AI video systems.

Lowering the barrier to AI video

Wan3.0 is now available for public testing through Alibaba Cloud’s Model Studio, the Wanxiang website and the Qwen creation platform on PC, with limited access also available through the Qwen app.

API pricing is set at 0.3 yuan, 0.6 yuan and 1.2 yuan per second for 480P, 720P and 1080P video generation respectively. Full API access will be opened soon.

The next AI video battleground

Wan3.0 reflects a broader shift in AI video generation — from experimental content creation toward practical production workflows.

The ability to generate 30-second videos gives AI models enough length to support more complete storytelling, while document-to-video conversion changes the starting point of creation from writing prompts to simply providing existing business materials.

The global AI video market is increasingly becoming a two-sided competition between U.S. and Chinese developers.

Companies including OpenAI and Google are advancing general-purpose video models, while Chinese players are focusing heavily on integration with business ecosystems and commercial deployment.

Wan3.0’s direct competitors in China include ByteDance’s Seedance and Kuaishou’s Kling, both of which have emerged as leading domestic AI video generation platforms.

With API prices falling to a fraction of previous levels, AI video generation is moving from an expensive creative experiment toward a more accessible productivity tool.

As a PPT file can now become a finished video, the next competition in AI video may not only be about who generates the most realistic images — but who makes creation start from the most useful inputs.