arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

代理化身与低秩缓存结合:实时单样本情感可控肖像动画

Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation

Haijie Yang, Jindi Bao, Yixuan Dong, Hongliang Zhang, Jian Bi, Hao Tang, Zhenyu Zhang, Jianjun Qian, Jian Yang

arXiv 2608.01978首次发表:更新:

AI 中文总结

该研究提出结合代理化身与低秩缓存的级联框架,解决音频驱动肖像动画的实时性与情感控制问题,实现了情感表达、身份保留及推理效率的提升。

AI 中文摘要

基于扩散生成模型的音频驱动肖像动画已快速发展,但兼具情感表达能力的实时单样本生成仍具挑战性。现有方法常存在情感感知运动先验不足、多步去噪过程中外观计算成本高昂的问题。为解决这些问题,我们提出Proxy Avatar Meets Low-Rank Caching,这是一个用于实时单样本情感可控肖像动画的级联框架。我们的方法不直接从音频生成目标肖像,而是使用基于高斯的情感代理化身作为可复用的运动生成器,该生成器在单个身份上训练一次,可从音频和情感标签生成富有表现力的驱动视频。由于代理化身仅提供运动而非目标外观或几何,一个大规模单样本重定向模型进一步从代理表现中提取与身份无关的运动,并将其适配到任意目标肖像。为提升推理效率,我们引入带有低秩缓存的零样本外观复用,该方法在初始去噪步骤缓存参考外观特征,并使用轻量级低秩适配器建模后续特征变化。大量实验表明,我们的方法实现了更强的情感表达、更好的身份保留动画,且推理成本大幅降低,可实现实时单样本肖像动画。

英文摘要

Audio-driven portrait animation has advanced rapidly with diffusion-based generative models, yet real-time one-shot generation with expressive emotion control remains challenging. Existing methods often suffer from insufficient emotion-aware motion priors and expensive appearance computation during multi-step denoising. To address these issues, we propose Proxy Avatar Meets Low-Rank Caching, a cascaded framework for real-time one-shot emotion-controllable portrait animation. Instead of directly generating the target portrait from audio, our method uses a Gaussian-based emotion proxy avatar as a reusable motion generator, which is trained once on a single identity to produce expressive driving videos from audio and emotion labels. Since the proxy avatar only provides motion rather than target appearance or geometry, a large-scale one-shot retargeting model further extracts identity-independent motion from the proxy performance and adapts it to arbitrary target portraits. To improve inference efficiency, we introduce zero-shot appearance reuse with low-rank caching, which caches reference appearance features at the initial denoising step and models subsequent feature variations using lightweight low-rank adapters. Extensive experiments demonstrate that our method achieves stronger emotional expressiveness, better identity-preserving animation, and substantially reduced inference cost, enabling real-time one-shot portrait animation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑