arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越唇形同步:基于参考的口腔细化用于音频驱动的人像动画

Beyond Lip Sync: Reference-Grounded Oral Refinement for Audio-Driven Portrait Animation

Bangxun Tang

arXiv 2609.38019首次发表:更新:

发表机构

University of California, Irvine(加利福尼亚大学尔湾分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有唇形同步系统渲染通用嘴部而非人物自身嘴部的问题,提出基于参考的口腔细化框架RGOR,利用注册帧和绕过VAE的高清嘴部补丁,配合配对判别器,在多数指标上达到最佳或次佳,保留人物唇齿细节。

AI 中文摘要

我们提出了RGOR(基于参考的口腔细化),一个音频驱动的唇形同步框架,它渲染的是被配音的具体人物的嘴部,而非通用的嘴部。现有的唇形同步系统紧密跟随音频并保持面部可识别,但它们渲染的嘴部是平均化的嘴部:嘴唇的形状和纹理、牙齿的排列以及嘴张开时露出的牙齿数量都不是该人物的。这个问题之所以持续存在,是因为当前的训练或评估中没有任何环节要求呈现人物自身的嘴部:感知损失接受任何合理的嘴部,面部身份主要由嘴部周围的皮肤承载,且修复系统的公开推理代码使用未掩码的目标帧作为参考,这掩盖了差距。为解决此问题,RGOR将每一生成帧都条件化于同一人物单独注册录制中的帧以及绕过VAE的嘴部高清补丁,并训练生成器对抗一个配对判别器,该判别器将每个渲染的嘴部与人物参考进行比较,并学会拒绝他人的逼真嘴部。我们进一步构建了一个评估协议,并用其比较开源和商业唇形同步系统在保留身份上的表现。实验表明,RGOR在大多数指标上达到最佳或次佳结果,并在保持同步和面部其余部分完整的同时,保留了人物自身的嘴唇和牙齿细节。

英文摘要

We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rather than a generic one. Existing lip-sync systems follow the audio closely and keep the face recognizable, yet the mouth they render is an average mouth: the shape and texture of the lips, the arrangement of the teeth, and how much of them shows as the mouth opens are not that person's. The problem persists because nothing in current training or evaluation asks for the person's own mouth: perceptual losses accept any plausible mouth, face identity is carried mostly by the skin around it, and the released inference code of inpainting systems uses the unmasked target frame as the reference, which hides the gap. To address this, RGOR conditions every generated frame on frames from separate enrollment recordings of the same person and on HD patches of the mouth that bypass the VAE, and trains the generator against a paired judge that compares each rendered mouth with the person's reference and learns to reject a realistic mouth of someone else. We further build an evaluation protocol and use it to compare open-source and commercial lip-sync systems on held-out identities. Experiments show that RGOR achieves the best or second-best result on most metrics, and preserves the person's own lip and dental detail while keeping synchronization and the rest of the face intact.

Comments19 pages, 8 figures, 5 tables. Under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑