arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

像你一样说话:实时说话头生成中模仿你的说话方式

Talk Like You: Imitating How You Speak in Real-Time Talking Head Generation

Baiqin Wang, Zhixing Ding, Jijie Li, Jiankuo Zhao, Zhen Lei, Xiangyu Zhu

arXiv 2610.06658首次发表:更新:

发表机构

MAIS, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; CAIR, HKISI, Chinese Academy of Sciences; SCSE, FIE, M.U.S.T.(中国科学院自动化研究所多模态人工智能系统全国重点实验室; 中国科学院大学人工智能学院; 中国科学院香港创新研究院人工智能与机器人创新中心; 澳门科技大学创新工程学院计算机科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出TalkLikeYou框架,通过运动空间建模和一步流匹配实现实时说话头生成,采用两阶段模仿学习捕捉个体说话习惯,并引入PLAD指标评估,显著提升模仿效果。

AI 中文摘要

在日常生活中,每个人都会表现出独特的说话习惯,即使在发同一个词时,也会导致细微但一致的唇形变化。尽管最近的说话头生成方法在视觉保真度和唇形同步方面取得了令人印象深刻的结果,但它们在很大程度上忽视了用户特定的定制,尤其是表征个体说话习惯的运动模式。这些习惯难以建模和捕捉,因为它们的运动模式高度细微,并且在不同个体之间往往相似。因此,许多方法会产生过于统一的面部运动,无法捕捉多样化的、因人而异的发音模式。为了解决这个问题,我们提出了TalkLikeYou,一个高效的框架,在说话头生成中模仿目标人物的说话方式。我们的方法在运动空间中建模习惯,并通过流匹配在推理时仅需一步采样即可实现实时性能。我们进一步采用两阶段模仿学习策略来捕捉习惯之间的细微差别,允许用户通过数据集中的预设风格或参考视频来指定目标习惯。此外,我们引入了一个新指标PLAD,将嘴部运动投影到代表性的发音轴上,以评估模仿准确性和生成多样性。大量实验表明,TalkLikeYou能够实时生成高质量的说话头,并且与先前方法相比,显著提高了说话习惯的模仿效果。代码可在以下网址获取:此https URL

英文摘要

In daily life, each person exhibits unique speaking habits, leading to subtle yet consistent lip-shape variations even when pronouncing the same word. Although recent talking head generation methods have achieved impressive visual fidelity and lip synchronization, they largely overlook user-specific customization, especially the motion patterns that characterize individual speaking habits. These habits are difficult to model and capture, as their motion patterns are highly fine-grained and often similar across individuals. As a result, many approaches produce overly uniform facial motions and fail to capture diverse, person-specific articulation patterns. To address this, we propose TalkLikeYou, an efficient framework that imitates how a target person speaks in talking head generation. Our method models habit in motion-space and achieves real-time performance through Flow Matching with only one sampling step during inference. We further adopt a two-stage imitation learning strategy to capture subtle distinctions between habits, allowing users to specify a target habit through either a preset style from the dataset or a reference video. In addition, we introduce a new metric PLAD that projects mouth motions onto representative articulation axes to evaluate imitation accuracy and generation diversity. Extensive experiments demonstrate that TalkLikeYou generates high-quality talking heads in real-time and significantly improves speaking habit imitation compared with prior methods. The code is available at: https://github.com/BQ-Wang0511/TalkLikeYou

Comments18 pages,10 figures. Project Page: https://bq-wang0511.github.io/TalkLikeYou/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑