Alibaba's Wan 3.0 has quietly dropped with 20 multimodal references and 30-second generations at ten cents a second
Wan 3.0 from Alibaba slipped out with almost no fanfare, bringing up to 20 multimodal reference inputs, 30-second video generation, and a pricing floor of $0.10 per second at 720p. Early testers say it belongs in the same conversation as MiniMax H3 and LTX 2.5.
Alibaba's video generation team shipped Wan 3.0 with almost no runway. No countdown. No launch event. One day it was not there; the next it was accepting up to 20 multimodal references in a single generation and producing clips up to 30 seconds long. The model landed while everyone was looking elsewhere, and that is almost certainly why it has been underreported.
The headline capability is the reference count. Twenty multimodal inputs means a filmmaker can feed the model a character image, a location still, a lighting reference, a costume detail, a prop sheet, and several mood frames simultaneously, all shaping one output. At the current ceiling of competing cloud models, the practical limit is typically two to four image references, occasionally six. Going to twenty is not a marginal increment; it is a different way of working, one that looks more like a traditional pre-production brief than a prompt box.
What is actually confirmed
Per creator @VictorInFocus, Wan 3.0 supports up to 20 multimodal references, generates up to 30 seconds of video per prompt, and prices 720p output at $0.10 per second. That puts a full 30-second generation at $3.00 at 720p, which is competitive with Seedance 2.5 at its standard rates and meaningfully cheaper at higher resolutions if Alibaba's pricing holds as resolution scales. The model is available via the Alibaba Wan platform now.
What has not been said
No higher resolution pricing has been published. It is not yet confirmed whether 1080p output is available, at what cost, or whether the 30-second limit applies at all resolutions. The precise mix of modalities allowed across the 20 reference slots, whether that means video clips, audio, depth maps or only images, has not been formally documented in a public spec sheet. These are not small gaps. A filmmaker planning a campaign around specific reference types needs to know before committing to a workflow.
No benchmarks against concurrent models have been published by Alibaba. The comparison to MiniMax H3, LTX 2.5 and others is currently word-of-mouth from early testers, @VictorInFocus among them, calling the first results genuinely competitive. Anecdotes are not benchmarks.
How it sits against the current field
At the moment Wan 3.0 arrived, working AI filmmakers already had Seedance 2.5 at 1080p, MiniMax H3 pushing toward local deployment, LTX 2.5 with open weights for real-time local generation, and FLUX 3 bringing native audio to Luma. The model market is genuinely crowded. What Wan 3.0 offers that none of those does is the 20-reference ceiling. If that number holds in practice and the quality is competitive, it solves a very specific pain point: visual consistency across a multi-element scene without relying on ControlNet stacks or iterative rerolling.
The $0.10/s pricing at 720p also deserves attention. For iterative testing, where a filmmaker might run 15 to 20 variations before locking a scene, a lower per-generation cost extends the creative runway without stretching the budget. Cloud pricing at this tier has typically been the argument for sticking with Seedance 2.0 over 2.5; Wan 3.0 positions itself as a third option at a similar price point with a different capability profile.
What to test first
The obvious first test is the reference stack. Build a simple five-element brief, character, location, prop, lighting reference, and mood frame, and measure how faithfully the model integrates all five without one overriding the others. That is where multi-reference systems tend to fail: the model latches onto the dominant visual and discards the rest. If Wan 3.0 genuinely balances 20 inputs, that is worth knowing immediately.
The second test is the 30-second generation. Most models degrade in temporal consistency after eight to ten seconds. Whether Wan 3.0 holds character identity and camera logic across a full 30-second clip is the question that determines whether it is useful for short-form narrative or only for assets and B-roll.
For AI filmmakers managing episodic or campaign work, Wan 3.0 is the first model in the current generation to make a credible pitch for the reference-heavy, pre-production-aligned workflow. Whether the output quality backs that pitch up is something only a proper test can confirm. Given the pricing, that test costs $3.00 to find out.
---
Sources: @VictorInFocus — first hands-on report and demo