arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29225cs.CV

ComplexSync:复杂场景下的高保真实时唇形同步

ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios

Jiaran Cai, Xingpei Ma, Shenneng Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对扩散模型在复杂场景下唇形同步质量差且推理慢的问题,提出ComplexSync统一框架,通过双流训练、蒸馏加速和关系对齐损失,实现超70 FPS的实时高保真唇形同步,并首个构建复杂场景基准。

中文摘要 AI 辅助

唇形同步旨在生成与语音音频精确对齐的视觉唇部动态。尽管扩散模型具有较高的生成质量,但它们在复杂场景中往往表现不佳,且推理延迟过高,限制了实际部署。我们提出了ComplexSync,一个统一的基于扩散的框架,能够在复杂条件下实现实时、高保真的唇形同步。首先,我们引入了一种双流联合训练策略,以减轻参考帧的信息泄漏,同时保持自然动态。其次,我们开发了一种基于蒸馏的加速方案,用于单步去噪,实现了超过70 FPS的吞吐量。第三,我们提出了一种关系对齐损失,利用视觉基础模型(VFMs)的结构先验来增强对复杂场景因素的鲁棒性。此外,我们提出了首个专门针对复杂唇形同步的基准,包含超过200个具有挑战性的视频序列和专门的评估指标。大量实验表明,ComplexSync在标准和复杂场景下均达到了最先进的性能,同时实现了实时推理。

英文摘要

Lip synchronization aims to generate visual lip dynamics that align precisely with speech audio. Despite the high generation quality of diffusion models, they often struggle in complex scenarios and suffer from prohibitive inference latency, limiting real-world deployment. We present ComplexSync, a unified diffusion-based framework that enables real-time, high-fidelity lip sync under complex conditions. First, we introduce a dual-stream joint training strategy to mitigate information leakage from reference frames while preserving natural dynamics. Second, we develop a distillation-based acceleration scheme for single-step denoising, achieving a throughput of over 70 FPS. Third, we propose a relational alignment loss that leverages structural priors from Vision Foundation Models (VFMs) to enhance robustness against complex scene factors. Furthermore, we present the first benchmark specifically designed for complex lip synchronization, comprising over 200 challenging video sequences and specialized metrics. Extensive experiments demonstrate that ComplexSync achieves state-of-the-art performance across both standard and complex scenarios while enabling real-time inference.

发表机构

  • Guangzhou Quwan Network Technology(广州趣丸网络科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑