EfficientSync:基于变形参考纹理混合的实时唇形同步
EfficientSync: Real-Time Lip Synchronization via Deformation-Based Reference Texture Mixing
查看机构详情
- The Hong Kong University of Science and Technology(香港科技大学)
- South China University of Technology(华南理工大学)
- University of Rochester(罗切斯特大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
EfficientSync是一种基于变形的实时唇形同步框架,通过动态纹理混合器等技术保留参考纹理,在HDTF和VFHQ数据集上以166 FPS实现领先的视觉质量与身份保留效果。
中文摘要 AI 辅助
音频驱动的唇形同步任务旨在操纵说话人脸视频的嘴部区域,使其与驱动音频匹配,同时保留头部姿态、身份和背景。尽管该任务本质上是局部编辑,但主流方法会使用基于GAN或扩散的重型解码器重建整个下半张脸,导致延迟显著,更关键的是会生成口腔内部细节(如牙齿和唇纹)的幻觉,而非保留真实纹理。我们认为身份保留的瓶颈并非参考帧稀缺,而是缺乏能忠实地传递参考帧中已有真实纹理的机制。因此,我们提出EfficientSync,这是一种基于变形的实时框架,用于保留参考纹理而非重新合成它们。首先,动态纹理混合器(Dynamic Texture Mixer)将多参考融合重新表述为通道级选择,在全局上下文中评估每个空间对齐的参考,并通过通道级加权求和聚合它们,以低成本保留纹理完整性。其次,时空移位自适应掩码(Spatio-Temporal Shifted Adaptive Masking)将源帧分解为嘴部生成条件和独立背景先验,抑制下半张脸泄漏,同时将合成的嘴部无缝融合到背景中。第三,STAR采样(STAR Sampling)作为零开销预处理步骤,检索最清晰且拓扑多样性最高的参考帧。在HDTF和VFHQ数据集上的实验表明,该方法在单个GPU上以166 FPS的帧率实现了领先的视觉质量和身份保留效果。视频演示可访问此https URL。
英文摘要
Audio-driven lip synchronization manipulates the mouth region of a talking-face video to match the driving audio while preserving head pose, identity, and background. Although the task is inherently local editing, prevailing approaches reconstruct the entire lower face with heavy GAN- or diffusion-based decoders, incurring substantial latency and, more critically, hallucinating intra-oral details such as teeth and lip wrinkles instead of preserving authentic textures. We contend that the bottleneck in identity preservation is not the scarcity of reference frames, but the lack of a mechanism that faithfully transfers the genuine textures they already contain. We therefore present EfficientSync, a real-time deformation-based framework that retains reference textures rather than resynthesizing them. First, the Dynamic Texture Mixer reformulates multi-reference fusion as channel-wise selection, evaluating each spatially aligned reference in a global context and aggregating them by channel-wise weighted summation, preserving textural integrity at low cost. Second, Spatio-Temporal Shifted Adaptive Masking decomposes the source frame into lip-generation conditions and an independent background prior, suppressing lower-face leakage while blending the synthesized mouth seamlessly into the background. Third, STAR Sampling, a zero-overhead pre-processing step, retrieves the sharpest and most topologically diverse reference frames. Experiments on HDTF and VFHQ show state-of-the-art visual quality and identity preservation at 166 FPS on a single GPU. Video demos: https://alunaticat.github.io/EfficientSync/index.html.