arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

E³mo-Bench:基于贝叶斯成对对齐的可扩展多模态诱发与表达情感理解基准

E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment

Lancheng Gao, Ziheng Jia, Shengyan Li, Zixuan Xing, Jiarui Wang, Huiyu Duan, Xiongkuo Min

arXiv 2608.10796首次发表:更新:

AI 中文总结

该研究针对现有多模态情感理解基准的不足,推出E³mo-Bench基准,提出贝叶斯成对对齐方法与E³mo-Score智能体,验证了框架有效性并指出MLLMs在情感任务中的缺陷,为多模态情感智能发展指明方向。

AI 中文摘要

理解表达情感与诱发情感对多模态大语言模型(MLLMs)实现全面的情感感知交互至关重要。然而,现有基准通常孤立考察表达与诱发情感,或受限于粗粒度、不完整的情感表征。为填补这一空白,我们推出E³mo-Bench,这是一个可扩展基准,包含2524个视频对应的12314个问答对,且具有预定义的情感视角。该基准通过三项互补任务评估诱发与表达情感理解:情感感知、开放词汇识别,以及效价-唤醒-支配性(VAD)评估。为高效扩展可靠的连续标注,我们提出贝叶斯成对对齐(Bayesian Pairwise Alignment),将稀疏、低负担的成对判断聚合为锚点参考的VAD估计值。此外,我们开发了E³mo-Score,这是一种无训练智能体,可聚合五模型委员会的互补判断以改进VAD估计。大量实验验证了我们框架的有效性,并揭示了诱发与表达情感范式间存在显著的性能偏差。这些发现,结合MLLMs在细粒度识别与维度评估方面的持续缺陷,为推进多模态情感智能指明了清晰方向。

英文摘要

Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interactions. However, existing benchmarks typically examine expressed and evoked emotions in isolation or are constrained to coarse-grained and incomplete affective characterizations. To bridge this gap, we introduce E$^3$mo-Bench, a scalable benchmark comprising $12{,}314$ question-answer pairs across $2{,}524$ videos with predefined affective perspectives. It evaluates evoked and expressed emotion understanding via $3$ complementary tasks: emotion perception, open-vocabulary recognition, and valence-arousal-dominance (VAD) assessment. To efficiently scale reliable continuous annotations, we propose Bayesian Pairwise Alignment, which aggregates sparse, low-burden pairwise judgments into anchor-referenced VAD estimates. Furthermore, we develop E$^3$mo-Score, a training-free agent that aggregates complementary judgments from a five-model committee to improve VAD estimation. Extensive experiments validate the effectiveness of our framework and expose a pronounced performance skew between evoked and expressed emotion paradigms. These findings, coupled with MLLMs' persistent deficits in fine-grained recognition and dimensional assessment, chart a clear course for advancing multimodal emotional intelligence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑