arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FMReward:基于人类偏好的音频驱动三维面部动画对齐与评估框架

FMReward: Aligning and Evaluating Audio-Driven 3D Facial Animation with Human Preferences

Sijing Wu, Yunhao Li, Zhilin Gao, Huiyu Duan, Yucheng Zhu, Guangtao Zhai, Patrick Le Callet

arXiv 2608.15296首次发表:更新:

发表机构

Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University; USC-SJTU Institute of Cultural and Creative Industry, Shanghai Jiao Tong University; Institut Universitaire de France (IUF); University of Nantes(上海交通大学图像通信与网络工程研究所; 上海交通大学 USC-文化创意产业学院; 法国大学研究院; 南特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文构建首个音频驱动三维面部动画人类偏好数据集FMPair,提出FMReward奖励模型与FMFL微调算法,可更好对齐人类偏好,提升该类动画的感知质量。

AI 中文摘要

音频驱动的三维面部动画对提升虚拟体验的沉浸感与交互性至关重要。尽管近期研究已展现出良好性能,但现有方法的训练与评估通常依赖基于真值的误差指标,难以与人类偏好对齐。为解决该问题,本文提出一套综合框架,从人类偏好数据中学习自动感知模型,并利用该模型改进和评估音频驱动三维面部动画的感知质量。首先,构建FMPair(面部运动成对偏好)——首个针对音频驱动三维面部动画的人类偏好数据集,该数据集通过系统化标注流程构建,包含来自8834段真实场景音频片段的65574个标注三维面部运动对。基于该成对比较数据集,本文提出面部运动奖励模型FMReward,其以音频和三维面部运动为输入,预测与人类偏好对齐的感知质量分数。在FMReward基础上,进一步提出面部运动奖励反馈学习(FMFL)算法,这是一种直接微调算法,利用预训练的奖励模型优化基于扩散模型的音频驱动三维面部动画,使其更好地对齐人类偏好。大量实验表明,FMReward在与人类偏好对齐方面优于其他指标,且FMFL在提升音频驱动三维面部动画的感知质量方面具有有效性。

英文摘要

Audio-driven 3D facial animation is essential for advancing immersion and interactivity in virtual experiences. Although recent advances have shown promising capabilities, the training and evaluation of existing methods typically rely on ground-truth-based errors, which fall short of aligning with human preferences. To address this, we present a comprehensive framework that learns an automatic perceptual model from human preference data and leverages it to improve and evaluate the perceptual quality of audio-driven 3D facial animation. To begin with, we construct FMPair (Facial Motion Pairwise preference), the first human preference dataset for audio-driven 3D facial animation, which is built through a systematic annotation pipeline and comprises 65,574 annotated 3D facial motion pairs from 8,834 distinct in-the-wild audio clips. Based on the pairwise comparison dataset, we propose a Facial Motion Reward model, termed FMReward, which takes audio and 3D facial motion as inputs and predicts a perceptual quality score aligned with human preferences. Building upon FMReward, we further introduce Facial Motion reward Feedback Learning (FMFL), a direct fine-tuning algorithm that leverages a pretrained reward model to optimize diffusion-based audio-driven 3D facial animation models for better alignment with human preferences. Extensive experiments demonstrate the superiority of FMReward over other metrics in aligning with human preferences and the effectiveness of FMFL in improving the perceptual quality of audio-driven 3D facial animation.

CommentsAccepted for publication in IEEE TVCG, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑