发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VibeAvatar通过解耦语音运动学适配器与美学运动策略,在轻量流模型中实现高保真、高效且美观的说话头像合成。
AI 中文摘要
多模态说话头像合成旨在从参考肖像和语音生成逼真的说话视频。尽管基于扩散的方法取得了快速进展,现有方法仍难以同时实现准确的唇部发音、人类偏好的运动美学和高效推理。我们观察到,语音准确性和运动美学源于根本不同的来源,应在互补阶段处理,而不是由单一生成器隐式学习。基于这一见解,我们提出VibeAvatar,通过语音运动学适配器(PKA)在条件阶段将面向识别的语音特征转换为语音运动学条件,以及美学运动策略(AMP)在后训练阶段通过组相对策略优化(GRPO)优化流一致的随机采样策略,从而解耦这两个目标。借助在紧凑的1D基于扭曲的潜在运动空间中运行的轻量级基于流的运动生成器,VibeAvatar在发音、美学和效率方面的客观指标和用户研究均达到最先进水平,同时仅需约3GB显存即可在10秒内生成10秒的512像素视频。
英文摘要
Multi-modal talking avatar synthesis aims to generate realistic talking videos from a reference portrait and speech. Despite rapid progress in diffusion-based methods, existing approaches still struggle to jointly achieve accurate lip articulation, human-preferred motion aesthetics, and efficient inference. We observe that phonetic accuracy and motion aesthetics arise from fundamentally different sources and should be addressed at complementary stages rather than learned implicitly by a single generator. Based on this insight, we propose VibeAvatar, which disentangles these two objectives through a Phonetic Kinematics Adapter (PKA) that converts recognition-oriented speech features into phonetic-kinematic conditions at the conditioning stage, and an Aesthetic Motion Policy (AMP) that optimizes a flow-consistent stochastic sampling policy via Group Relative Policy Optimization (GRPO) at the post-training stage. With a lightweight flow-based motion generator operating in a compact 1D warp-based latent motion space, VibeAvatar achieves state-of-the-art results in articulation, aesthetics, and efficiency on both objective metrics and user studies, while generating a 10-second 512px video in under 10 seconds with only $\sim$3GB VRAM.