arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2502.10841cs.CV

SkyReels-A1:视频扩散Transformer中的富有表现力的肖像动画

SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers

  • Cranberry-Lemon University(蔓越莓柠檬大学)

机构由 AI 辅助整理,请以论文原文为准。

Di Qiu, Zhengcong Fei, Rui Wang, Jialin Bai, Changqian Yu, Mingyuan Fan, Guibin Chen, Xiang Wen

更新

AI总结:

针对现有肖像动画方法存在的身份失真、背景不稳定等问题,提出基于视频扩散Transformer的SkyReels-A1框架,通过表情感知条件模块、面部图像-文本对齐模块及多阶段训练,提升运动迁移精度、身份保留与时间连贯性,适用于虚拟化身等领域。

AI中文摘要:

我们提出SkyReels-A1,这是一个基于视频扩散Transformer构建的简单而有效的框架,用于促进肖像图像动画。现有方法仍面临身份失真、背景不稳定和不真实面部动态等问题,尤其是在仅头部动画场景中。此外,扩展以适应不同身体比例通常会导致视觉不一致或不自然的关节运动。为解决这些挑战,SkyReels-A1利用视频DiT强大的生成能力,提高面部运动迁移精度、身份保留和时间连贯性。该系统集成了一个表情感知条件模块,能够通过表情引导的关键点输入驱动无缝视频合成。整合面部图像-文本对齐模块加强了面部属性与运动轨迹的融合,增强了身份保留。此外,SkyReels-A1采用多阶段训练范式,在确保稳定身份重现的同时逐步细化表情与运动之间的相关性。大量实证评估表明,该模型能够产生视觉连贯且构图多样的结果,使其在虚拟化身、远程通信和数字媒体生成等领域具有高度适用性。

英文摘要:

We present SkyReels-A1, a simple yet effective framework built upon video diffusion Transformer to facilitate portrait image animation. Existing methodologies still encounter issues, including identity distortion, background instability, and unrealistic facial dynamics, particularly in head-only animation scenarios. Besides, extending to accommodate diverse body proportions usually leads to visual inconsistencies or unnatural articulations. To address these challenges, SkyReels-A1 capitalizes on the strong generative capabilities of video DiT, enhancing facial motion transfer precision, identity retention, and temporal coherence. The system incorporates an expression-aware conditioning module that enables seamless video synthesis driven by expression-guided landmark inputs. Integrating the facial image-text alignment module strengthens the fusion of facial attributes with motion trajectories, reinforcing identity preservation. Additionally, SkyReels-A1 incorporates a multi-stage training paradigm to incrementally refine the correlation between expressions and motion while ensuring stable identity reproduction. Extensive empirical evaluations highlight the model's ability to produce visually coherent and compositionally diverse results, making it highly applicable to domains such as virtual avatars, remote communication, and digital media generation.

↑