发表机构
University of Technology Sydney; Shandong University of Science and Technology(悉尼科技大学; 山东科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出首个跨身份舌部动态迁移框架,结合 foundation model 辅助的自举流水线与空间约束潜在掩码扩散模型,在舌部指标上较基线提升超两倍,且 VLM 评估证实其感知优势。
AI 中文摘要
现代人脸重演系统采用几何驱动表示实现了令人印象深刻的姿态和表情迁移,但大多忽略舌部动态,导致语音和表情动作期间口腔内部在解剖结构上不一致。本文提出了首个用于人脸重演的跨身份舌部动态迁移框架,引入 foundation model 辅助的自举流水线,无需精心标注即可生成适用于野外重演的专用舌部分割模型;还提出空间约束的潜在掩码扩散模型用于真实舌部合成,采用自适应掩码膨胀实现口腔边界的平滑过渡。大量实验表明,在所有舌部专用指标上,该方法较所有基线提升超过两倍;此外,本文提出基于 VLM 的评估协议可大规模复制专家标注,证实所有消融变体均具有感知优势。
英文摘要
Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tongue dynamics, leading to anatomically inconsistent mouth interiors during speech and expressive motions. We introduce the first framework for cross-identity tongue dynamics transfer in face reenactment. We propose a foundation-model-assisted bootstrapping pipeline that produces a dedicated tongue segmentation model for in-the-wild reenactment without curated annotations. We further introduce a spatially constrained latent masked diffusion model for realistic tongue synthesis, with adaptive mask dilation for seamless mouth boundary transitions. Extensive experiments demonstrate improvements of more than two times over all baselines on every tongue-specific metric. We additionally propose a VLM-based evaluation protocol that replicates expert annotation at scale, confirming perceptual superiority across all ablation variants.