让我看着你:用于对话语音合成的高级面部表情建模
Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
浏览论文内容
中文总结 AI 辅助
研究针对对话语音合成中面部表情线索常被忽视及缺乏多模态数据集的问题,提出基于大语言模型的FacialTalker框架,含AUTokenizer和DualDPO策略,构建VSDD-1K数据集,实验表明该框架在表情感知和语音合成质量上表现出色。
中文摘要 AI 辅助
对话语音合成是人机交互的基本组成部分,旨在生成上下文合适、富有表现力和同理心的语音。然而,面部表情编码了对同理心语音交互至关重要的微妙而丰富的情感线索,而现有方法往往忽略了这一重要模态。此外,缺乏大规模的语音和视觉模态自然对话数据集也限制了对话中视觉情感理解的发展。为了解决这些限制,我们提出了FacialTalker,一个基于大语言模型主干构建的面部表情感知CSS框架。为了有效地编码面部表情,我们提出了AUTokenizer,一种单码本视觉tokenizer,它将每个帧级面部表情离散化为一个紧凑的token,通过面部动作单元组合的监督进行训练。我们进一步引入了双直接偏好优化(DualDPO)策略,该策略通过对视觉和语音token序列联合施加偏好约束来扩展DPO,以增强模型在多模态对话上下文中对面部表情和语音语义的理解。此外,我们构建了VSDD-1K,一个通过全自动管道从真实世界互联网对话中收集的大规模多模态对话数据集,包括超过1033小时的同步说话者视频和语音,超过85%的帧包含有效面部。广泛的客观和主观实验表明,FacialTalker在面部表情感知和语音合成质量方面始终优于强大的基线,生成的语音更自然、更有表现力,并且与对话上下文更好地对齐。结果也验证了我们的训练策略和数据集构建管道的有效性。
英文摘要
Conversational Speech Synthesis is a fundamental component of human-computer interaction, aiming to generate contextually appropriate, expressive, and empathetic speech. However, facial expressions encode subtle and rich affective cues that are crucial for empathetic speech interaction, whereas existing approaches often overlook this important modality. In addition, the lack of large-scale natural conversational datasets with both speech and visual modalities also limits the development of visual affect understanding in conversational settings.To address these limitations, we propose FacialTalker, a facial-expression-aware CSS framework built upon a large language model backbone. To efficiently encode facial expressions, we propose AUTokenizer, a single-codebook visual tokenizer that discretizes each frame-level facial expression into a compact token, trained with supervision from combinations of facial Action Units. We further introduce a dual direct preference optimization (DualDPO) strategy, which extends the DPO by jointly imposing preference constraints on both visual and speech token sequences, to enhance the model's understanding of facial expressions and speech semantics in multimodal conversational contexts. Moreover, we construct VSDD-1K, a large-scale multimodal dialogue dataset collected through a fully automated pipeline from real-world Internet conversations, comprising over 1,033 hours of synchronized speaker videos and speech, with more than 85\% of frames containing valid faces. Extensive objective and subjective experiments demonstrate that FacialTalker consistently outperforms strong baselines in facial-expression perception and speech synthesis quality, generating speech that is more natural, expressive, and better aligned with the conversational context. The results also validate the effectiveness of our training strategy and dataset construction pipeline.
发表机构
- Inner Mongolia University(内蒙古大学)
机构由 AI 辅助整理,请以论文原文为准。