DualTrack:通过预训练先验的对称耦合实现同步语音-手势生成
DualTrack: Synchronized speech-gesture generation via symmetric coupling of pretrained priors
浏览论文内容
中文总结 AI 辅助
DualTrack通过对称耦合预训练先验,在12.5 Hz时间线上协调语音和手势生成,在BEAT2多语言评估中优于GELINA基线。
中文摘要 AI 辅助
联合语音-手势合成必须在有限配对数据下协调两种模态。现有方法往往缺乏双向交互、语言覆盖有限,或简化了身体和手指的表示。我们提出DualTrack,它在共享的12.5 Hz时间线上耦合预训练的语音和运动先验。因果适配器交换先前数据包的信息,而当前状态融合在流分别完成十六码本数据包之前协调流。我们评估了四种语言的43个BEAT2录音,说话者从联合训练和验证中排除。在共享的英语/西班牙语输入上,在没有语音或运动前缀的情况下,DualTrack在词错误率和全运动Fréchet手势距离上取得更低值,在节拍一致性和语音自然度上优于评估的GELINA基线。
英文摘要
Joint speech-gesture synthesis must coordinate two modalities despite limited paired data. Existing approaches often lack bidirectional interaction, have limited language coverage, or simplify body and finger representations. We present DualTrack, which couples pretrained speech and motion priors on a shared 12.5 Hz timeline. Causal adapters exchange previous-packet information, while current-state fusion coordinates the streams before they separately complete sixteen-codebook packets. We evaluate 43 BEAT2 recordings in four languages, with speakers held out from joint training and validation. On the shared English/Spanish inputs, without speech or motion prefixes, DualTrack achieves lower word error rate and full-motion Fréchet Gesture Distance, higher beat consistency and speech naturalness than the evaluated GELINA baseline.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- X Square Robot
机构由 AI 辅助整理,请以论文原文为准。