arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PAIR:实时多模态对话智能体中的感知情感推断与调节

PAIR: Perceptual Affective Inference and Regulation in a Real-Time Multimodal Conversational Agent

Kexin Quan, Zijian Ding, Jiaye Yong, Qinshi Zhang, Dong Wang, Jessie Chin

arXiv 2610.07523首次发表:更新:

发表机构

School of Information Sciences, University of Illinois at Urbana-Champaign; School of Computing, KAIST; Department of Computer Science, Florida State University(伊利诺伊大学厄巴纳-香槟分校信息科学学院; 韩国科学技术院计算学院; 佛罗里达州立大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出PAIR实时多模态对话智能体,通过评估脚手架推断情感并调节,14天部署验证了推断准确性与支持模式的相关性。

AI 中文摘要

持续的情感支持要求生成式智能体将瞬时的情感推断和调节与跨会话的连续性联系起来。我们提出了PAIR(感知情感推断与调节),一个实时多模态智能体,它重构事件如何被评估为情感状态。评估脚手架产生效价-唤醒-支配度估计并选择调节指导,通过协调的语音、颜色和虚拟形象线索的对话来传递。滚动记忆跨会话携带上下文,脚手架在指导后重新运行。在为期14天的部署中,有19名参与者,1093次会话将初始和指导后的估计与无锚定的自我报告配对。初始效价在9点SAM量表上达到MAE 1.20(r=.68),支配度达到MAE 1.30,唤醒度即使在粗化后也显示出较弱的一致性。自我报告的情感变化随初始状态而变化,在从负效价开始的会话中效价增幅最大。感知理解与更大的效价增加相关,并且与数值预测误差几乎没有对应关系。在两周内,帮助性增加而输入缩短;访谈将个性化和陪伴追溯到相关回忆、上下文更新和熟悉的对话。这些发现通过逐事件、第一人称评估将推断准确性与支持的对话和时间模式联系起来。

英文摘要

Sustained emotional support requires generative agents to connect momentary emotion inference and regulation with continuity across encounters. We present PAIR (Perceptual Affective Inference and Regulation), a real-time multimodal agent that reconstructs how an event is appraised into an emotional state. Appraisal scaffolds produce a valence-arousal-dominance estimate and select regulation guidance, delivered through conversation with coordinated speech, color, and avatar cues. Rolling memory carries context across sessions, and the scaffold re-runs after guidance. In a 14-day deployment with 19 participants, 1,093 sessions paired initial and post-guidance estimates with unanchored self-reports. Initial valence reached MAE 1.20 on the 9-point SAM scale (r=.68), dominance reached MAE 1.30, and arousal showed weak agreement even after coarsening. Self-reported emotional change varied with initial state, with the largest valence increases in sessions that began at negative valence. Perceived understanding was associated with greater valence increase and showed little correspondence with numerical prediction error. Over two weeks, helpfulness increased while input shortened; interviews traced personalization and companionship to relevant recall, context updates, and familiar dialogue. These findings connect inference accuracy to conversational and temporal patterns of support through per-event, first-person evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑