arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06009cs.CV

Wan-Animate-2:拓展角色动画的应用边界

Wan-Animate-2: Pushing the Application Boundaries of Character Animation

  • Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室)

机构由 AI 辅助整理,请以论文原文为准。

Guangyuan Wang, Li Hu, Dechao Meng, Zhongyi Zhang, Peng Zhang, Xindi Zhang, Mingyang Huang, Ruoshi Zhang, Ke Sun, Zhe Zhang, Xingjun Wang, Gang Cheng, Hai Xu, Bang Zhang

中文总结 AI 辅助

本文提出Wan-Animate-2角色动画框架,通过消除中间运动提取器提升动画保真度,新增文本驱动视角控制,推出实时高效变体Wan-Animate-2-Lite,将发布其基础模型权重。

中文摘要 AI 辅助

角色图像动画仍是计算机视觉领域一项基础但极具挑战性的任务。现有方法大致可分为三类范式:基于显式运动表示的方法存在提取误差和身份漂移问题;基于隐式运动特征的方法因压缩而丢失细粒度动态信息;基于上下文学习的方法虽避免了中间表示,但会产生过高的计算成本。此外,当前所有系统均为离线合成设计,无法满足数字化身、直播主播等交互式应用的实时性要求。为解决这些局限,本文提出Wan-Animate-2,这是一款端到端的角色动画框架,可在重新设计的扩散Transformer中直接使用驱动视频。我们的架构通过完全消除中间运动提取器,实现了更优的运动保真度和身份保留能力。我们还引入了文本驱动的视角控制,将输出相机视角与驱动视频解耦——这一能力在依赖显式运动表示的现有角色动画方法中很少得到支持。除生成质量外,本文还提出了Wan-Animate-2-Lite,这是一款高效变体,通过三阶段训练范式将推理延迟降至实时阈值:带误差缓冲机制的教师强制预训练,以及分块反向传播的自强制蒸馏。这使得交互式应用的流式角色动画成为可能,开辟了此前无法实现的新部署场景。定性评估和用户研究表明,Wan-Animate-2在不同角色和运动模式下均能实现高保真动画结果。为推动进一步研究和社区发展,我们将向公众发布Wan-Animate-2-Base模型权重。

英文摘要

Character image animation remains a foundational yet challenging task in computer vision. Existing approaches can be broadly categorized into three paradigms: methods based on explicit motion representations suffer from extraction errors and identity drift; methods based on implicit motion features lose fine-grained dynamics through compression; and in-context learning approaches avoid intermediate representations but incur prohibitive computational costs. Furthermore, all current systems are designed for offline synthesis, unable to meet the real-time requirements of interactive applications such as digital avatars and live-streaming hosts. To address these limitations, we present Wan-Animate-2, an end-to-end character animation framework that directly consumes the driving video within a redesigned Diffusion Transformer. Our architecture achieves superior motion fidelity and identity preservation by eliminating intermediate motion extractors entirely. We further introduce text driven viewpoint control that decouples the output camera perspective from the driving video--a capability rarely supported by prior character animation methods that rely on explicit motion representations. Beyond generation quality, we present Wan-Animate-2-Lite, an efficient variant that reduces inference latency to real-time thresholds through a three-stage training paradigm: teacher forcing pretraining with error buffer mechanism, and Self-Forcing distillation with chunk-wise backpropagation. This enables streaming character animation for interactive applications, opening new deployment scenarios that were previously infeasible. Qualitative evaluations and user studies demonstrate that Wan-Animate-2 achieves high-fidelity animation results across diverse characters and motion patterns. To foster further research and community development, we will release the Wan-Animate-2-Base model weights to the public.

补充信息

相关深度报道

↑