arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16011cs.GRcs.CVcs.LGcs.MMcs.ROcs.SDeess.AS

EMODY Flow:情感感知的音频驱动全身动作生成

EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

  • INRIA(法国国家信息与自动化研究所)
  • Univ. Grenoble Alpes(格勒诺布尔阿尔卑斯大学)
  • CNRS(法国国家科学研究中心)
  • Grenoble INP(格勒诺布尔国立理工学院)
  • GIPSA-lab(GIPSA实验室)

机构由 AI 辅助整理,请以论文原文为准。

Harsh Kumar Agarwal, Xavier Alameda-Pineda, Olivier Perrotin

AI总结:

EMODY Flow提出轻量级流匹配框架,利用冻结的Qwen-3 Omni和Mimi音频编解码器,通过辅助情感分类器恢复情感敏感性,生成情感感知的全身动作,在BEAT2上显著提升手势质量并零样本迁移到面部动画。

AI中文摘要:

具身对话智能体需要与语音和情感状态同步的全身动作(身体姿态和面部表情)。全模态大语言模型在多模态理解方面表现出色,但仅产生语言输出,在具身响应生成方面存在关键空白。我们识别并解决了一个情感条件化失败的问题:与其他弱化弱条件信号的生成器一样,当流匹配模型同时获得丰富的音频嵌入和离散的情感标签时,会抑制情感,产生几乎相同的动作,而不管指定的情感如何。我们提出了EMODY Flow,一个轻量级(约35M参数)的流匹配框架,它附加到冻结的Qwen-3 Omni模型上,并重用其内部的Mimi音频编解码器来条件化两个并行的DiT生成器——一个用于SMPL-X身体姿态,一个用于FLAME面部表情。训练时的辅助情感分类器通过强制生成的动作具有情感可识别性来恢复情感敏感性。EMODY Flow在BEAT2手势质量上树立了新的最先进水平,FGD为0.302,节拍相关性为0.853,多样性为24.62——分别比之前的最佳结果提高了26%、5%和62%——并且无需特定领域的微调即可零样本迁移到TFHP上的面部动画。除了这些定量提升外,分类器还产生了明显情感分离的动作,我们通过对生成手势的多维缩放分析进行了定性展示。

英文摘要:

Embodied conversational agents require synchronized full-body motion (body gestures and facial expressions) that aligns with speech and emotional state. Omni-modal large language models excel at multimodal understanding but produce only linguistic outputs, leaving a critical gap in embodied response generation. We identify and address a failure of emotion conditioning: like other conditional generators that under-use weak conditioning signals, a flow-matching model given both a rich audio embedding and a discrete emotion label suppresses the emotion, generating near-identical motion regardless of the specified emotion. We present EMODY Flow, a lightweight (around 35M parameters) flow-matching framework that attaches to a frozen Qwen-3 Omni model and reuses its internal Mimi audio-codecs to condition two parallel DiT generators - one for SMPL-X body pose, one for FLAME facial expressions. A training-time auxiliary emotion classifier restores emotion sensitivity by forcing generated motion to be emotion-identifiable. EMODY Flow sets a new state of the art on BEAT2 gesture quality, with FGD 0.302, Beat Correlation 0.853, and Diversity 24.62 - improving over the best prior results by 26%, 5%, and 62% respectively - and transfers to zero-shot facial animation on TFHP without domain-specific fine-tuning. Beyond these quantitative gains, the classifier yields clearly emotion-separated motion, which we demonstrate qualitatively through a multidimensional-scaling analysis of the generated gestures.

补充信息

↑