arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EmpaAva:开源智能体式3D化身共情实时聊天机器人

EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot

Jie Yang, Wenhao Xu, Shuhui Lin, Hao Fei

arXiv 2608.04709首次发表:更新:

发表机构

National University of Singapore; Tsinghua University; University of Oxford(新加坡国立大学; 清华大学; 牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出首个开源智能体式3D化身共情聊天机器人EmpaAva,通过三智能体架构实现闭环多模态交互,在评估中优于多款基准模型,将开源并提供在线演示。

AI 中文摘要

本文提出了EmpaAva,据我们所知,它是首个开源的智能体式3D化身共情聊天机器人,将共情响应生成(ERG)从纯文本交互拓展至实时面对面互动。通过类视频通话界面,用户与3D数字人对话,该数字人可从语音及可选视觉信息中感知用户情感,并用带情感的语音、唇形同步的面部动作以及照片级真实感的3D高斯渲染进行回复。其核心是大语言模型(LLM)协调的三智能体架构,其中感知、共情响应规划与具身渲染形成闭环,搭配响应规划层将每个回复编译为可执行的多模态计划,使语音、表情与渲染保持同一共情意图。EmpaAva基于成熟的开源模块构建,提供将这些模块绑定为可控、可检查体验的智能能力。在自动与人工评估中,EmpaAva在情感理解、响应质量及视听一致性上均优于纯文本、2D聊天脸及多模态化身基准模型,我们将开源EmpaAva并提供在线实时演示。

英文摘要

This paper presents EmpaAva, to our knowledge the first open-source, agentic 3D-avatar empathetic chatbot, which carries empathetic response generation (ERG) from text-only exchanges into live, face-to-face interaction. Through a video-call-like interface, a user speaks to a 3D digital human that reads their affect from speech and optional vision, and replies with emotional speech, lip-synced facial motion, and photorealistic 3D Gaussian rendering. At its core, an LLM coordinates a Tri-Agent Architecture, in which perception, empathetic response planning, and embodied rendering form a closed loop, paired with a Response Planning layer that compiles each reply into an executable multimodal plan, keeping voice, expression, and rendering on one empathetic intent. Building on strong open-source modules, EmpaAva supplies the intelligence that binds them into one controllable, inspectable experience. In automatic and human evaluations, EmpaAva surpasses text-only, 2D talking-face, and multimodal avatar baselines in emotion understanding, response quality, and audio-visual consistency. We open-source EmpaAva with an online live demo.

CommentsProject&Demo: https://empaava.top/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑