发表机构
Hainan International College, Communication University of China; School of Data Science and Intelligent Media, Communication University of China; School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中国传媒大学海南国际学院; 中国传媒大学数据科学与智能媒体学院; 中山大学深圳校区网络空间安全学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对共语手势生成中语义利用不足的问题,提出可靠性感知的语义节奏控制框架,通过双分支信息增益估计与噪声条件调制,实现语义表现力与节奏同步的平衡。
AI 中文摘要
共语手势生成旨在合成既与语音在时间上同步又与口语内容在语义上一致的自然手势。尽管近期方法能够生成节奏上合理的手势动作,但它们往往过度依赖声学韵律,而未能充分利用文本语义,尤其是在语义标注不完整、含噪声或不可用的情况下。因此,生成的手势可能跟随语音节奏,却未能表达预期的语义。为解决这一问题,我们提出了一种用于共语手势生成的可靠性感知语义节奏控制框架。我们首先学习一个离散动作先验,在紧凑且结构化的动作编码空间中表示连续手势。随后,我们引入一种双分支语义贡献估计机制,由全多模态分支和仅音频分支组成。它们之间的分布差异被表述为条件信息增益,以量化文本语义在多大程度上改变预测动作。基于这一估计,一个可控的语义节奏目标选择性地增强内容相关片段中的语义引导,同时限制节奏主导片段中不必要的语义干预。此外,我们将背景噪声视为声学可靠性条件,并引入噪声条件特征调制以及有益的潜在扰动,以提升在现实声学环境下的生成鲁棒性。在基准数据集上的实验表明,所提出的框架在语义表现力、节奏同步性、动作多样性和鲁棒性之间取得了良好的平衡,实现了可靠且可控的共语手势生成。
英文摘要
Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible motions, they often rely heavily on acoustic prosody while underutilizing textual semantics, especially when semantic annotations are incomplete, noisy, or unavailable. Consequently, the generated gestures may follow speech rhythm while failing to express the intended semantics. To address this problem, we propose a reliability-aware semantic-rhythm control framework for co-speech gesture generation. We first learn a discrete motion prior that represents continuous gestures in a compact and structured motion-code space. We then introduce a dual-branch semantic contribution estimation mechanism consisting of a full multimodal branch and an audio-only branch. Their distributional discrepancy is formulated as conditional information gain to quantify how much textual semantics changes the predicted motion. Based on this estimate, a controllable semantic-rhythm objective selectively strengthens semantic guidance in content-relevant segments while limiting unnecessary semantic intervention in rhythm-dominant segments. Furthermore, we treat background noise as an acoustic reliability condition and introduce noise-conditioned feature modulation together with beneficial latent perturbation to improve generation robustness under realistic acoustic environments. Experiments on benchmark datasets demonstrate that the proposed framework achieves a favorable balance among semantic expressiveness, rhythmic synchronization, motion diversity, and robustness, enabling reliable and controllable co-speech gesture generation.
Comments9 pages, 5 figures, 3 tables