arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WanPE:面向现代文生视频的电影级提示增强

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

arXiv 2609.30221首次发表:更新:

发表机构

Nanjing University; Wan Team, Alibaba Group; Fudan University; Tsinghua University(南京大学; 阿里巴巴集团万相团队; 复旦大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

WanPE是一个397B参数的提示增强模型,通过视频锚定的反向构建和SC-GRPO实现导演级电影规划,在5-15秒内大幅提升人类偏好,并在30秒领域与Seedance 2.5竞争。

AI 中文摘要

视频生成始于文本空间,通过编写电影剧本,然后物化为像素。随着当代视频生成器扩展到30秒并忠实遵循复杂条件,文本提示在很大程度上引导了制作过程,规划动作、摄像机轨迹、灯光和声音如何在多镜头序列中展开。在本文中,我们提出了WanPE,一个397B参数的提示增强模型,在1.05M个真实世界视频上训练,以掌握导演级别的电影规划。WanPE通过视频锚定的反向构建来制定镜头级别的电影计划,并采用语义一致性GRPO(SC-GRPO)来在镜头间和时间上忠实保留用户需求。为了基准测试这一能力,我们策划了WanPEval,一个人工标注的测试平台,覆盖5至30秒的时长和不同的意图粒度,并辅以约11K次盲法成对评估。当为Wan3.0的视频生成器提供动力时,WanPE-397B在5-15秒内将人类偏好相对于原始用户提示提升了10.66-18.84个百分点,在30秒领域则大幅提升了50.86个百分点。消融研究表明,反向构建在优于前向重写方面表现出明显优势,而SC-GRPO在不同模型规模下稳健地保持了语义保真度。最终,WanPE在5-15秒内领先所有评估的商业产品,并在30秒时与Seedance 2.5保持竞争力。

英文摘要

Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by approximately 11K blind pairwise assessments. When powering Wan3.0's video generator, WanPE-397B boosts human preference over raw user prompts by 10.66-18.84 points at 5-15 seconds and by a dramatic 50.86 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑