发表机构
National University of Singapore; Southern University of Science and Technology; Fudan University; Harbin Institute of Technology (Shenzhen)(新加坡国立大学; 南方科技大学; 复旦大学; 哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态共情回复生成中情绪冲突、缺乏引导和错误传播问题,提出EmAvatar框架,通过冲突消解和迭代细化实现精确感知与表现力合成,在四个任务上超越现有方法。
AI 中文摘要
基于虚拟形象的多模态共情回复生成已成为以人为中心的系统中的关键能力,旨在识别用户情绪并合成带有同步文本、音频和说话人脸视频的回复。尽管近期取得了进展,现有方法仍存在三个关键局限:(1)忽视跨模态的情绪冲突,(2)缺乏明确的多模态合成引导,(3)忽视多模态回复生成固有的错误传播。为解决这些局限,我们提出EmAvatar,一个用于精确情绪感知和表现力回复生成的新框架。它首先通过暴露模态间预测冲突来进行深思熟虑的多模态情绪识别,然后在冲突检查器与证据收集器之间启动多轮问答过程以收集证据进行冲突消解,从而产生稳健的、基于证据的预测。在回复生成方面,EmAvatar首先合成一个复合脚本,将文本回复与表现力指令耦合。此外,为确保高质量合成,一种迭代细化机制评估并修订脚本直至其符合预定标准,为后续音频和视频合成提供可靠引导。跨四个任务的广泛实验表明,EmAvatar优于最先进的方法。我们的代码将公开发布。
英文摘要
Avatar-based multimodal empathetic response generation has emerged as a pivotal capability in human-centric systems, aiming to recognize user emotions and synthesize responses with synchronized text, audio, and talking-face video. Despite recent progress, existing methods still suffer from three critical limitations: (1) overlooking conflicting emotions across modalities, (2) lacking explicit multimodal synthesis guidance, and (3) neglecting inherent error propagation of multimodal response generation. To address these limitations, we propose EmAvatar, a novel framework for precise emotion perception and expressive response generation. It first performs deliberative multimodal emotion recognition by exposing inter-modal prediction conflicts and then initiates a multi-round QA process between a Conflict Inspector and an Evidence Collector to gather evidence for conflict resolution, leading to a robust, evidence-aware prediction. Regarding response generation, EmAvatar first synthesizes a composite script that couples the textual response with an expressive instruction. Moreover, to ensure high-quality synthesis, an iterative refinement mechanism evaluates and revises the script until it aligns with predefined criteria, serving as reliable guidance for subsequent audio and video synthesis. Extensive experiments across four tasks demonstrate that EmAvatar outperforms state-of-the-art methods. Our code will be publicly released.