发表机构
University of Maryland, College Park; Netflix(马里兰大学帕克分校; 奈飞)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ATR,一种多语言唇形同步评判器,通过先对齐后推理的方法,在七语言基准上显著提升配音质量评估的AUC,并泛化至多种LLM和下游任务。
AI 中文摘要
配音质量控制需要一个无参考的评判器,它能够仅利用无声视频和文本,判断候选文本行是否在内容和时序上与说话者的可见口型相匹配,因为配音音频可能尚不存在。现有的视觉语音识别器和视频-语言模型并不适合这一场景:即使经过微调以从唇部运动恢复口语内容,它们对时序错误仍然在很大程度上不敏感。我们引入了“先对齐后推理”(ATR),一种多语言唇形同步评判器,它首先建立帧级唇部表示与候选行的语音单元之间的单调对齐,然后基于该对齐进行推理以做出最终判断。一个对齐评分器为LLM提供每个语音单元的局部证据和一个校准后的全局对齐分数,使其能够联合推理内容和时序。在七语言基准上,我们的方法相对于相应的Qwen3.5 SFT基线,在2B、4B和9B推理器上的平均AUC分别提升了59.4%、50.2%和50.8%。这些提升在不同LLM家族中具有泛化性,在LLaMA-3.1-8B和Mistral-7B上相对于最佳基线的平均AUC分别提升了45.9%和46.6%。它们还跨数据集迁移到三个未见过的MuAViC语言。此外,我们在两个基于真实配音行构建的下游任务上进行了评估。在配音行重排序中,ATR-9B比最佳唇读基线高出52.0%,而在脚本到片段分配中,ATR-9B比最佳唇读基线高出17.7%。
英文摘要
Dubbing quality control requires a reference-free judge that can determine whether a candidate text line matches a speaker's visible articulation in both content and timing, using only silent video and text because dubbed audio may not yet exist. Existing visual speech recognizers and video-language models are poorly suited to this setting: even when fine-tuned to recover spoken content from lip motion, they remain largely insensitive to temporal errors. We introduce $\textit{Align Then Reason}$ (ATR), a multilingual lip-sync judge that first establishes a monotonic alignment between frame-level lip representations and the phonetic units of the candidate line, then reasons over this alignment to make the final judgment. An alignment scorer provides the LLM with both local evidence for each phonetic unit and a calibrated global alignment score, enabling it to reason jointly about content and timing. On a seven-language benchmark, our method improves mean AUC over the corresponding Qwen3.5 SFT baselines by 59.4%, 50.2%, and 50.8% with 2B, 4B, and 9B reasoners, respectively. The gains generalize across LLM families, reaching mean AUC improvements of 45.9% and 46.6% over the best baseline for LLaMA-3.1-8B and Mistral-7B, respectively. They also transfer across datasets to three unseen MuAViC languages. Furthermore, we evaluate on two downstream tasks built from real dubbing lines. On dub-line reranking, ATR-9B outperforms the best lip-reading baseline by 52.0%, while on script-to-clip assignment, ATR-9B improves over the best lip-reading baseline by 17.7%.