AI Video Generation
Alibaba's Wan3.0 sells video by the second: $6 for a 30-second clip
Wan3.0 can turn a PDF or a brand deck into a 30-second video in one generation, with per-second API pricing that tops out at $0.20 for 1080P. We break down the per-clip math and the gaps Alibaba admits in its own testing.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-08-16 · 4 min read

Alibaba's Wan3.0, the latest generation of the Wan video model family, is now live on Alibaba Cloud Model Studio with per-second billing: $0.05 at 480P, $0.10 at 720P, $0.20 at 1080P. A full 30-second generation runs $6 at the top tier, and $1.50 as a 480P draft. The numbers matter less than what they enable: iterating on cheap drafts, then paying up for the final pass.
Most AI video models produce clips of a few seconds, enough for a shot but not a story, as Alibaba puts it. Wan3.0 extends single-generation length to 30 seconds, with a smart duration feature that suggests a length from the prompt and a video extension option for continuing past the initial output. Longer output changes the creative vocabulary: room for narrative pacing, continuous camera moves, and one-take sequences. Long clips are a known failure point across the video AI stack, and Nvidia built AV-Flamingo for long-form understanding because most models fall apart on longer footage.
The per-second meter is where the economics turn concrete. At 480P, a 30-second draft costs $1.50, cheap enough for rough iteration. Finishing the same length at 1080P costs $6 per attempt. Alibaba's own example runs the math the same way: iterate at 480P, finish at 1080P. Fixed per-second billing looks calm next to the token-pricing churn elsewhere in AI: a reply to Zhipu's viral GLM-5.2 launch claimed the price had already 10x'd.
| Resolution | Price per second | 30-second clip |
|---|---|---|
| 480P | $0.05 | $1.50 |
| 720P | $0.10 | $3.00 |
| 1080P | $0.20 | $6.00 |
Two things stay unresolved. The listing does not say whether failed or discarded runs still get billed, and it names four creative modes (reference, editing, replication, driving) without defining what separates them. It also caps concurrency at two tasks, which means a redo queues behind whatever is already running. For teams chasing character consistency, the meter at least makes each miss predictable: $6 per do-over.
The bigger functional change is input. Wan3.0 is the first model in the Wan family to read documents as structured input: doc, xls, ppt, pdf, txt, key, pages, numbers, and md, up to one file or link per request, capped at 100MB and 50 pages. Pair a document with a prompt, and the prompt acts as the control center that interprets the input and shapes the output. Alibaba's demo is a café brand introduction document plus one sentence of direction, "Turn this brand story into a warm-toned brand TVC," producing a complete brand film.
Alibaba's framing is that video generation is shifting from content novelty to production tooling. The everything-to-video capability targets advertising, UI interaction demos, and destination marketing, where destination films no longer require large-scale location shoots. In Alibaba's reading, the same path covers software feature animations and data visualization: the model preserves the aesthetics and texture of the original design while elevating it into motion. Nothing in the API decides whether those use cases pay off, but the input flexibility is what lets an existing deck become a draft video without rebuilding the idea from scratch. Alibaba is not the only one making the everything-in-one-model bet; MiniMax's H3 unifies text, image, video, and audio generation in one model.
On output quality, Wan3.0 leans on two claims: lifelike, diverse faces instead of the glossy, interchangeable look of typical AI characters, and precise reference-to-video consistency across characters, props, spaces, and style, including distinct emotions in group scenes. The editing capability introduced in Wan2.7 carries forward, so visuals, plot, and dialogue can be modified without regenerating from scratch. Iteration becomes a revision pass, not a do-over.
Consistency is the metric production teams care about most, the blog argues: in reference-to-video mode, Wan3.0 reproduces reference details precisely and holds them stable across key dimensions. That is the difference between a clip that looks right in isolation and one that survives a cut sequence.
Alibaba also flags where the model falls short. In internal testing, audio texture and on-screen text rendering accuracy "are improving but not yet where we want them," and both are active refinement areas. For a model that positions documents and charts as input, legible on-screen text is a nontrivial gap. Wan3.0 is not the only video model with limits. A new arXiv study found Gemini 3.6 Flash counts state changes but misses blinks.
The trajectory is plain. The Wan family has shipped eight iterations from Wan1.0 to Wan3.0, and Alibaba aims those same capabilities, understanding richer information and generating more physically consistent output, at embodied AI, autonomous driving, and industrial simulation. Wan is not Alibaba's only model track; the company recently open-sourced its most capable model, Qwen 3.8-Max. For now, the practical read is simpler. The meter makes AI video spend predictable. Whether $6 buys a usable clip is a question only a production run can answer.
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.