发表机构
University of Southern California; Cantina Inc.(南加州大学; Cantina公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PVSync是一个统一唇形同步模型,通过窗口级对比学习和音素级发音目标实现音视频偏移估计与发音评分,在多个基准上优于现有方法。
AI 中文摘要
嘴唇运动可以与语音的时序匹配,但不一定匹配发出的声音。我们引入PVSync,一个用于音视频偏移估计和音素级发音评分的统一模型。PVSync结合了窗口级对比学习用于同步,以及一个音素级发音目标,该目标在片段间对齐同一视素类别的音频和视频嵌入。视素将具有相似可见发音的音素分组。视素标签自动从强制对齐的转录中推导,无需手动标注。在来自13个说话头视频生成模型的偏移校正视频上,PVSync的唇形同步质量排名比LSE-C更接近人类排名,实现了0.83对0.34的Spearman相关性。在从保留语音自动构建的基准上,PVSync区分视素匹配与不匹配的音视频对的ROC AUC为0.91。PVSync在保留的真实世界片段上的时间偏移恢复方面也优于SyncNet和MTD-VocaLiST。代码和基准数据将在录用后发布。
英文摘要
Lip movements can match the timing of speech without matching the spoken sounds. We introduce PVSync, a unified model for audio-visual offset estimation and phoneme-level articulation scoring. PVSync combines window-level contrastive learning for synchronisation with a phoneme-level articulation objective that aligns audio and video embeddings of the same viseme class across clips. Visemes group phonemes with similar visible articulation. Viseme labels are derived automatically from forced-aligned transcripts, without manual annotations. On offset-corrected videos from 13 talking-head video generation models, PVSync matches human rankings of lip-sync quality more closely than LSE-C, achieving a Spearman correlation of 0.83 versus 0.34. On an automatically constructed benchmark from held-out speech, PVSync distinguishes viseme-matched from mismatched audio-visual pairs with an ROC AUC of 0.91. PVSync also outperforms SyncNet and MTD-VocaLiST in temporal offset recovery on held-out in-the-wild clips. Code and benchmark data will be released upon acceptance.