arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UltraDub:通过统一视觉引导的流学习与轨迹指导实现真实配音

UltraDub: Towards Authentic Dubbing by Unifying Visually-Steered Flow Learning and Trajectory Guidance

Gaoxiang Cong, Liang Li, Jianwei Wen, Zhedong Zhang, Zheng-Jun Zha, Qingming Huang

arXiv 2610.05932首次发表:更新:

发表机构

Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Hangzhou Dianzi University; ByteDance; University of Science and Technology of China(中国科学院计算技术研究所; 中国科学院大学; 杭州电子科技大学; 字节跳动; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

UltraDub提出统一视觉引导流学习与轨迹指导的配音框架,通过MDR模块和RTG机制提升唇形同步与语言准确性,在四个数据集上达到最先进性能。

AI 中文摘要

视觉语音克隆需要与可见发音同步的清晰、说话者一致的语音。然而,顺序多模态条件化可能破坏先前建立的时序和说话者线索,而不平衡的推理指导可以提高语言准确性,但以牺牲唇形同步为代价。在本文中,我们提出UltraDub,一个统一视觉引导的流学习与轨迹指导的配音框架,它以两种方式利用视觉:作为多模态上下文聚合的连续运动,以及作为轨迹校正的结构节奏。具体来说,我们引入了运动引导的双上下文检索(MDR)模块,该模块通过共享的唇部运动查询残差持续重新校准语言和说话者风格检索,利用独立的时间条件门控来调节它们的贡献。此外,我们提出了节奏锚定的轨迹指导(RTG),一种无需训练的机制,在仅视觉的预测中点评估层次化多模态校正,安全地加强语义条件化,同时更好地保留时间对齐。最后,我们构建了DiverseDub,一个多场景基准,用于评估野外视频配音。大量实验表明,UltraDub在四个数据集上实现了最先进的性能。

英文摘要

Visual voice cloning requires intelligible, speaker-consistent speech synchronized with visible articulation. However, sequential multimodal conditioning can disrupt previously established temporal and speaker cues, while imbalanced inference guidance can improve linguistic accuracy at the expense of lip synchronization. In this paper, we propose UltraDub, a Unifying Visually-Steered Flow learning and trajectory Guidance Dubbing framework that leverages vision in two ways: as continuous motion for multimodal context aggregation, and as structural rhythm for trajectory rectification. Specifically, we introduce the Motion-guided Dual-context Retrieving (MDR) module, which continually recalibrates linguistic and speaker-style retrieval through shared lip-motion query residuals, utilizing independent time-conditioned gates to regulate their contributions. Furthermore, we propose Rhythm-anchored Trajectory Guidance (RTG), a training-free mechanism that evaluates hierarchical multimodal corrections at a visual-only predictive midpoint, safely strengthening semantic conditioning while better preserving temporal alignment. Finally, we construct DiverseDub, a multi-scenario benchmark to evaluate video dubbing in the wild. Extensive experiments demonstrate that UltraDub achieves state-of-the-art performance across four datasets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑