Xemo-Talker:为音频驱动的说话人像合成显式解锁情感
Xemo-Talker: Unlock Emotions Explicitly for Audio-Driven Talking Portrait Synthesis
AI总结:
针对音频驱动说话人像的情感控制难题,提出Xemo-Talker模型,通过非主成分监督的情感分支与Tri-Loss损失,实现了情感分类精度、唇形同步及推理效率的良好平衡。
AI中文摘要:
音频驱动的说话人像中的精确情感控制仍是一项挑战,因为现有系统依赖隐式情感调节,常导致控制间接且不足。此外,由于准确唇形同步与细粒度情感控制之间存在固有权衡,在整个运动空间中使用显式情感相关损失进行训练存在显著困难。本文揭示了一项关键发现:尽管情感线索分布在整个运动空间中,但将判别式监督集中在非主成分上可实现更好的情感-唇形同步平衡,因为主成分主要编码高能量的发音和姿态变化。基于这一见解,我们提出Xemo-Talker,该模型首先学习用于稳定发音和唇形同步的中性语音-运动映射,然后引入由非主子空间监督引导的轻量情感分支。为增强情感控制,我们设计了Tri-Loss,由类间分离、类内紧凑性和非主成分对比学习构成。给定音频输入、参考图像和情感标签,Xemo-Talker在保持有竞争力的唇形同步和高推理效率的同时,实现了最先进的情感分类准确率,性能接近真实数据的测量结果。源代码可在指定网址公开获取。
英文摘要:
Precise emotion control in audio-driven talking heads remains a challenge due to the reliance on implicit emotion regulation in existing systems, which often leads to indirect and insufficient control. Additionally, training with explicit emotion-related losses across the entire motion space poses significant difficulties due to the inherent trade-off between accurate lip synchronization and fine-grained emotion control. In this paper, we reveal a key finding: although emotional cues are distributed throughout the motion space, concentrating discriminative supervision on less-principal components achieves a better emotion-lip synchronization balance, as principal components mainly encode high-energy articulation and pose variations. Building on this insight, we propose Xemo-Talker, which first learns a neutral speech-to-motion mapping for stable articulation and lip synchronization, and then introduces a lightweight emotion branch guided by less-principal subspace supervision. To enhance emotion control, we design a Tri-Loss consisting of inter-class separation, intra-class compactness, and less-principal contrastive learning. Given an audio input, a reference image, and an emotion label, Xemo-Talker achieves state-of-the-art emotion classification accuracy while maintaining competitive lip synchronization and high inference efficiency, with performance approaching that measured on real videos.The source code is publicly available at https://github.com/chaolongy/Xemo-Talker.