SevenTnewS

AI Video Generation

Wan3.0-Video charges by the second: a 30-second clip runs to $6

Alibaba's Wan3.0-Video bills per second on DashScope, with a 30-second 1080P clip costing $6 per generation. We break down the pricing tiers, the multi-input workflow, and the questions the listing leaves unanswered.

Emmanuel Fabrice Omgbwa Yasse AI-assisted

2026-08-07 · 3 min read

Wan3.0-Video charges by the second: a 30-second clip runs to $6
Sources : Perceptron Mk1 …

Wan3.0-Video, Alibaba's video generation model on DashScope, bills per second of output. At 1080P, the top tier, each second costs $0.20. It generates clips up to 30 seconds long, so one full-length generation comes to $6 per attempt.

The lower tiers soften the math. 480P bills at $0.05 a second, or $1.50 for a 30-second clip. 720P sits between at $0.10 a second, or $3.00 for the same length. The bill for any clip is length times rate, which makes budgeting about as simple as video generation pricing gets.

Four creative modes, one model

Wan3.0-Video is listed as an all-in-one model rather than a plain text-to-video engine. The page names four creative capabilities: reference, editing, replication, and driving. It does not define those terms in any detail, so how "driving" differs from "reference" in practice is left to the user to discover. That is the single biggest gap in an otherwise concrete listing.

Graphique : Cost of a 30-second Wan3.0-Video clip by resolution
Costs at each resolution tier as listed in the article's rate card.

What is clear is the input side. The model accepts audio, image, text, and video, and it can parse files, web pages, and complex images as reference material. Paired with a claim of production-grade character consistency and lifelike visuals with sound, the pitch targets teams that need to reuse or control an established look, not teams firing off one-off prompts.

The per-second rate card

ResolutionPrice per secondCost for 30 seconds
480P$0.05$1.50
720P$0.10$3.00
1080P$0.20$6.00

The API details matter for anyone wiring this into a pipeline. Generations run asynchronously through DashScope, with two concurrent tasks, an async queue of 50, and 30 requests per minute. The reference implementation submits a prompt with a resolution, an adaptive ratio, and a duration, then polls for the finished clip. That profile fits batch jobs where you queue several generations and check back, rather than workflows that need synchronous turnaround.

Sora, Veo, Runway, and a crowded field

Wan3.0-Video enters a market that already has a long roster. Sora 2 generates video with synced audio, though its original API is deprecated and scheduled to shut down on September 24, 2026. Google offers Veo 3.1 and the cheaper Veo 3.1 Lite. Runway positions Gen-4.5 on high visual fidelity and creative control, and Luma Ray, Kling, and Seedance fill out the list. The Wan listing compares itself to none of them and gives no rival prices, so the per-second rates float without direct reference points. The closest public pricing signal comes from MiniMax, which prices output from its all-in-one H3 model at a third of mainstream per-second rates and ranks first in editing, per our MiniMax H3 coverage.

Alibaba already covers the other half of the video stack. Qwen3-VL is an open-weight video-language family with dense variants from 2B to 32B and mixture-of-experts variants up to 235B-A22B, with a native 256K token context. Wan3.0-Video on DashScope is the generation counterpart. The listing does not link the two explicitly, but teams already running Qwen3-VL for video understanding now have a generation option on the same platform. The open-weight push extends beyond video: Qwen3.8-Max, Alibaba's first open-weight Max-class model, landed on Hugging Face and ModelScope, per our report on the Qwen3.8-Max release.

What the listing leaves open

Three things are still unclear. The page never explains what the "driving" and "replication" modes do, and it shows no examples. It also does not say whether failed or discarded generations still get billed. And the two-task concurrency limit is hard to judge from a listing: redo one bad clip and the next generation queues behind whatever is already running.

None of that changes the headline number. Alibaba is selling seconds of video at $0.20 each at the top tier, and producing a character-consistent 30-second clip that misses the mark is a $6 do-over. Per-second billing at least makes that spend predictable. The output, as always, is the part that is not. That unpredictability is the gap a few projects are attacking: VideoCoCo scripts scenes in Blender code and lets a simulator play them out, a shift covered in our VideoCoCo explainer, while PhiZero teaches a video model a compact physical language before it renders, per the PhiZero write-up.

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.