Alibaba's Wan team has released WanPE, a 397-billion-parameter model that rewrites your short prompt into a full cinematic screenplay before video generation even begins
WanPE is a dedicated prompt enhancement model from the team behind Wan video generation. Feed it a brief description and it returns a director-level, shot-by-shot screenplay. Independent testing cited by HuggingPapers shows a preference improvement of up to 50.9 points over unenhanced prompts at 30 seconds of generated footage.
If you have spent any serious time with AI video generation, you already know the dirty secret of the craft: the model is rarely the bottleneck. The prompt is. Getting a system like Wan to render believable motion, consistent framing and a coherent sense of scene requires a level of written direction that most people simply do not have time to write from scratch, and that very few tutorials actually teach. Alibaba's Wan team appears to have decided to solve that problem not by tweaking the video model itself, but by building an entirely separate model whose sole job is to write better prompts than you would.
WanPE is a 397-billion-parameter prompt enhancement model, per @HuggingPapers. The scale is not incidental. At that parameter count it sits comfortably in frontier LLM territory, which is a meaningful signal about how seriously the Wan team is treating prompt quality as a first-class problem. Feed it a brief user description and it returns what is described as a director-level, shot-by-shot cinematic screenplay. The stated preference improvement is up to 50.9 points on outputs generated at the 30-second mark, compared to the same prompts entered without enhancement. Those figures come from the research paper rather than independent third-party evaluation, so treat them as directional rather than definitive until reproducible benchmarks appear.
What this changes in practice
The existing workaround for most AI filmmakers is manual prompt engineering: writing multi-paragraph scene descriptions that specify framing, lighting, duration, character action, camera movement and tone. That works, but it is slow, it requires a particular kind of writing discipline, and the results are inconsistent between users. A dedicated enhancement model that has been trained specifically on cinematic language sits above any general-purpose LLM you might use for the same task, because its entire parameter budget is pointed at one narrow output type.
The practical pipeline shift is significant. Instead of spending 20 minutes writing a detailed prompt and hoping the phrasing lands correctly, a filmmaker submits a logline-length description, WanPE expands it into structured shot language, and that structured output goes to Wan for rendering. The loop gets shorter. Iteration gets faster. And crucially, the variance between a skilled prompt engineer and a less experienced one should narrow, because the enhancement model is doing the translation work.
What was not said
Several things remain unclear from the current reporting. No release date for public access has been announced. It is not yet known whether WanPE will be integrated directly into Wan's consumer and API surfaces, offered as a standalone model call, or made available only through Alibaba's own platforms. The 50.9-point preference improvement figure comes from the paper's own evaluation methodology, and the specific benchmark design matters enormously for interpreting that number. A 50-point swing on a human preference study with a narrow evaluator pool is a different claim from a 50-point swing on a broad, diverse test set. No pricing has been disclosed. There is also no clarity yet on whether the model handles languages other than English or Chinese, which matters for the substantial non-English-speaking creator community that uses Wan.
How it sits against the alternatives
Several platforms already offer some form of prompt assistance, from simple auto-enhance toggles to Claude or GPT-4 integrations that expand descriptions on the fly. What distinguishes WanPE in principle is specificity: it is a frontier-scale model trained entirely for this task, not a general LLM being repurposed. Whether that specificity actually outperforms a well-prompted general model in practice is the first thing worth testing when access becomes available.
For filmmakers currently using Wan 3.0 specifically, this is a development worth watching closely. If WanPE ships as an integrated pre-processing step, the effective quality ceiling of Wan output rises without any change to the underlying video model. That is a more efficient upgrade path than waiting for a new model generation.
What to test first
When access does open, the immediate test is straightforward: take ten prompts you have already used successfully in Wan 3.0, run them through WanPE, and compare the enhanced versions against your originals. Pay attention to whether the model introduces camera movement language and shot transitions you would not have written yourself, and whether those additions survive into the final render or get ignored by the video model. The second test is failure modes: what happens when you give WanPE a genuinely ambiguous or abstract prompt. Enhancement models can overcorrect, adding cinematic specificity that was never intended and locking out creative interpretations the video model might otherwise have found.
For now, WanPE is a research announcement with compelling headline numbers and several important unknowns still to resolve. It is the kind of development that could quietly become one of the most-used tools in an AI filmmaker's chain, provided the access model is sensible and the preference gains hold outside the paper's own test conditions.