Wan 3.0 is now generally available, and its native 30-second single-pass generation changes what AI filmmakers can do without stitching

Alibaba's Wan 3.0 has launched across its own cloud platforms and more than a dozen third-party integrations simultaneously, bringing native 30-second clip generation, omni-reference inputs and reality-grade rendering to the tools filmmakers already use.

By Leeby Shmeeby

Wan 3.0 is out, it is generally available right now, and Alibaba has coordinated one of the broader simultaneous platform rollouts this year. The headline capability is native 30-second single-pass generation: you feed the model your inputs and it returns half a minute of video in one go, with no segment stitching and no seam-hiding required. If you have ever spent an afternoon wrestling two or three five-second clips into a coherent scene, you understand immediately why that matters.

The other central claim is omni-reference. Wan 3.0 accepts text, images, audio, video, web pages and documents as input in a single generation request. That is not a bullet-point feature list assembled for a press release; it is a meaningful shift in how you can brief a model. A mood board, a reference clip, a line of narration and a written scene description can all go in together, and the model is built to reconcile them into one output. Whether that reconciliation holds under adversarial conditions, mixing contradictory visual references, is something the community will stress-test in the coming days.

Wan 3.0 is now generally available, and its native 30-second single-pass generation changes what AI filmmakers can do without stitching

What was actually announced

Alibaba launched Wan 3.0 on 24 August 2026 via its own @Alibaba_Wan account, simultaneously announcing availability on Alibaba Cloud Model Studio and Qwen Cloud. A 30 percent launch discount on the Standard tier runs from 23 August to 23 September. The four headline capabilities named in official posts are: native 30-second generation, omni-reference input, reality-grade rendering with improved real-world motion, and dynamic motion handling. Researcher @haoyuc98 from the Wan model team attributed the visual quality gains to work with aesthetic experts and creative consultants, and to richer pre-training datasets across visual domains.

Where you can use it right now

The simultaneous platform availability is unusual in scale. Confirmed integrations announced on launch day include @fal, @pika_labs, @Scenario_gg, @magnific, @Segmind_ai via PixelFlow, @RunningHub_ai, @AskVenice, @Deevid_AI, @TapNow_AI, @LeonardoAi, @Flovaai, @envato, @Pixmax_ai, @runware, @nadou_pro, and @SeaArt_Ai. Pika is positioning its API Club pricing as up to 35 percent cheaper than competitors for Wan 3.0 access.

Pricing on Alibaba's own platforms

Standard API pricing at full rate, before the launch discount, is $0.05 per second at 480p, $0.10 per second at 720p, and $0.20 per second at 1080p. At those rates, a native 30-second 1080p generation costs $6.00 before any discount, or $4.20 during the promotional period. Those figures are worth holding in your head when comparing across the platforms above, since third-party pricing will vary. No date has been announced for when the launch discount ends beyond the stated 23 September cutoff.

What has not been said

Alibaba has not published a technical paper or benchmark comparisons against Seedance 2.5, MiniMax H3 or other current leaders. The omni-reference claim, specifically multi-format inputs in a single pass, is compelling but has not been independently verified at scale. The 30-second native generation length is a ceiling, not a floor; it is not yet clear whether the model degrades on shorter clips optimised for social formats. Temporal consistency across camera cuts within a 30-second run is a separate question from motion quality within a single shot, and community testing has barely begun. Quality at 480p versus 1080p has not been characterised publicly.

What to test first

For working AI filmmakers the most practical first experiment is the omni-reference path. Feed it a style-reference image, a brief audio clip, and a text scene description simultaneously, and see whether the output actually synthesises all three or simply prioritises one. That is the scenario where Wan 3.0 either earns its positioning or falls back to behaviour indistinguishable from a standard image-to-video model. Second, run the full 30 seconds on a scene requiring sustained character consistency across movement, since single-pass length is only useful if the model holds identity through the whole duration. The fal integration offers text-to-video, image-to-video and reference-to-video as separate endpoints, which makes isolating each capability straightforward for comparison.

The Pika partnership also deserves attention. Pika's API Club pricing claim of up to 35 percent below competitors, combined with Pika's existing audio and speech pipeline, opens up a workflow where Wan 3.0 handles the visual generation and Pika's own models handle voice and sound, all within one billing relationship. That may be the most immediately practical path for anyone who does not want to stitch together accounts across multiple platforms.