Alibaba Wan3.0 turns slide decks into 30-second videos

Wan3.0 hit public beta with native 30-second clips, document-to-video inputs, and pricing from $0.05 per second. Here is what changes for marketing teams and why Chinese labs keep stretching video length.

SaifullahSaifullah
4 min read
Alibaba Wan3.0 turns slide decks into 30-second videos

Most AI video tools still feel like GIF factories. You get three to five seconds, stitch clips in an editor, and pray the character's face does not morph between cuts. Alibaba Wan3.0 is betting the next fight is length and inputs, not prettier thumbnails.

The model entered public beta in early August 2026 through Alibaba Cloud Model Studio and Qwen Cloud. The headline specs: 30-second clips in one pass, document inputs (PPT, PDF, DOC, XLS), and API pricing from $0.05 per second at 480p.

I build lead sites and content pipelines for service businesses. Slide-to-video is not a toy feature for them. It is how a clinic turns a staff training deck into a patient-facing explainer without hiring a video crew.

What Wan3.0 actually ships

Wan3.0 is the successor to Wan2.7-Video. The jump from 15 to 30 seconds matters because it crosses from "social clip" to "single continuous scene with camera movement."

SpecWan3.0Prior generation (Wan2.7)
Max duration30 seconds (single pass)15 seconds
Input modalitiesText, image, audio, video, documents, webpagesText, image, audio, video
Document formatsdoc, xls, ppt, pdf, txt, key, pages, numbers, mdNot supported
Output resolutions480p, 720p, 1080pUp to 1080p
AudioNative audio-visual generationPrior versions improved audio
API pricing (480p)$0.05/sHigher on prior tiers

Alibaba's launch blog calls the document input feature Omni-Reference. You can feed a PowerPoint deck or PDF (max 100MB, 50 pages) and the model builds video from the static content. That is the use case I keep thinking about for local businesses: turn a service menu PDF into a short promo without a videographer.

Comparison of Wan3.0 specs including 30-second duration, document inputs, and pricing tiers

Pricing against Western competitors

The pricing story is aggressive. At 1080p ($0.20/s), a full 30-second clip costs $6. Google's Veo 3.1 standard tier runs about $0.40 per second according to The Next Web's coverage. That is a 2x gap at the top resolution tier.

ResolutionWan3.0 ($/second)30s clip cost
480p$0.05$1.50
720p$0.10$3.00
1080p$0.20$6.00

Chinese labs have been quietly stretching video length while Western attention stayed on chatbots. Wan3.0 landed the same week Alibaba raised roughly $10.2 billion in a Hong Kong share placement earmarked for AI capex. The sequencing is deliberate.

Access and rollout

This is not a download-and-run open weights release. Access is staged:

  • Public beta through Model Studio and Qwen Cloud (application required)
  • API preview with approval before you can call wan3.0-video
  • wan.video consumer site promised as members-only

The async API pattern uses DashScope endpoints with reference media arrays for character consistency. You can lock a character with reference video plus reference images and generate up to 30 seconds while keeping fur, clothing, and staging stable. Alibaba's demo page shows a tabby cat and wolf duo sparring for a full 30 seconds at 1080p.

Flow diagram from PowerPoint or PDF document through Wan3.0 to 30-second video output

Where I would use it (and where I would not)

Good fits for client work:

  • Training deck to patient explainer video for clinics
  • Product catalog PDF to short social clips for e-commerce
  • Real estate listing sheets to property walkthrough-style promos (with clear AI disclosure)

Bad fits:

  • Anything requiring legal precision without human review (medical claims, financial advice)
  • Brand campaigns where a single wrong frame destroys trust
  • Workflows that need frame-level editing control (Wan3.0 edits visuals and dialogue without full regeneration, but you still lack a traditional NLE timeline)

Seedance 2.5 launched around the same window with 30-second single-pass clips too. The differentiator for Wan3.0 is document-to-video. Seedance accepts more reference inputs (up to 50 across images, video, and audio). Neither has independent public benchmarks yet. Both are API-only closed models.

The bigger picture for Applied AI shipping

Video generation is moving from "cool demo" to "content ops layer." When a model accepts your existing PDFs and slides as input, the integration surface shifts from prompt engineering to document pipeline wiring.

That is the work I actually bill for: ingest the asset, validate the output, add human review gates, publish to the CMS, and track which clips convert. The model is one node in a graph, not the whole product.

If you are evaluating Wan3.0 against Veo or Runway for a content pipeline, test three things:

  1. Character consistency across the full 30 seconds on your actual brand assets
  2. Document fidelity when your slides have dense text (not just hero images)
  3. Total cost per published minute including retries and human QA time

Wan3.0 makes the economics interesting. It does not remove the need for editorial judgment.

Building a document-to-video pipeline for your marketing stack? Book a free discovery call and we can scope the integration.

Share this post

Related posts