arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11918cs.MMcs.AIcs.HC

从表面到深度:迈向多模态情感理解中的认知评估推理

From Surface to Depth: Towards Cognitive Appraisal Reasoning in Multimodal Emotion Understanding

Jia Li, Yichao He, Yangchen Yu, Qiankun Li, Xinyi Li, Baiyi Ye, Zhenzhen Hu, Richang Hong, Erik Cambria

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多模态情感理解的表面化问题,提出从感知到认知评估的新范式,构建了数据集、模型与基准,实现了更可靠的情感理解与更强的跨域泛化能力。

中文摘要 AI 辅助

近期,多模态大语言模型(MLLMs)越来越多地将可解释推理融入情感理解任务中。然而,主要基于可观测情感线索的推理会将情感理解简化为表面的线索-标签关联,从而产生“聪明汉斯效应”。当情感线索隐含存在、跨模态冲突、具有误导性或被冗余细节掩盖时,这类捷径方法的可靠性会大幅下降。相比之下,人类的情感是由个体对周围事件的解读与评估方式塑造的,而非仅基于可观测线索。受情绪评估理论启发,我们将多模态情感理解定义为从感知到认知评估的递进过程,并引入了支持这一新范式的数据集、模型和基准。CogEmo-40K是一个大规模指令调优数据集,通过从感知到评估的流程构建,可引出情感背后六个认知评估维度的证据支撑推理。CogEmo-MoE是一个紧凑的稀疏多模态大语言模型(MLLM),引入了交错的混合专家(MoE)模块以实现评估特定适配,其规模远小于典型的情感多模态大语言模型,却能实现有效的评估推理。CogEmo-Bench引入了评估证据质量分数(AEQS),用于评估六个互补评估维度下的认知-情感理解,解决了传统情感指标仅评估预测的情感是什么、不评估其产生原因的局限性。大量实验表明,我们的范式不仅在CogEmo-Bench上取得领先性能,还展现出强大的跨域泛化能力。研究结果表明,从感知到评估的推理可超越表面的线索-标签关联,实现更可靠的多模态情感理解,并使多模态大语言模型与人类的认知更趋一致。

英文摘要

Recent multimodal large language models (MLLMs) increasingly incorporate explainable reasoning for emotion understanding. However, reasoning based mainly on observable affective cues can reduce emotion understanding to superficial cue-label associations, giving rise to the Clever Hans effect. Such shortcuts become unreliable when affective cues are implicit, conflicting across modalities, linguistically misleading, or obscured by redundant details. In contrast, human emotions are shaped by how individuals interpret and evaluate surrounding events beyond observable cues. Inspired by appraisal theories of emotion, we formulate multimodal emotion understanding as a progression from perception to cognitive appraisal, and introduce a dataset, a model, and a benchmark to support this novel paradigm. CogEmo-40K is a large-scale instruction-tuning dataset constructed through a perception-to-appraisal pipeline to elicit evidence-grounded reasoning across six cognitive appraisal dimensions underlying emotion. CogEmo-MoE is a compact sparse MLLM that introduces interleaved MoE blocks for appraisal-specific adaptation, enabling effective appraisal reasoning at a substantially smaller scale than typical emotion MLLMs. CogEmo-Bench introduces an Appraisal Evidence Quality Score (AEQS) to assess cognitive-affective understanding across six complementary appraisal dimensions, addressing the limitation of conventional emotion metrics that evaluate what emotion is predicted but not why it arises. Extensive experiments show that our paradigm not only leads CogEmo-Bench, but also exhibits strong cross-domain generalization. Our findings suggest that perception-to-appraisal reasoning can move beyond surface-level cue-label associations toward more reliable multimodal emotion understanding and closer cognitive alignment between MLLMs and humans.

发表机构

  • Hefei University of Technology(合肥工业大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑