发表机构
Tsinghua University; Nankai University; Beijing Institute of Technology; Chinese Academy of Sciences(清华大学; 南开大学; 北京理工大学; 中国科学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多模态大语言模型视觉情商评估缺失问题,引入情感陈述判断公式ESJ,开发INSETS及MVEI基准,构建EmObserver模型。通过实验评估,确立了ESJ、MVEI和EmObserver在推进面向MLLM的视觉情商方面的作用。
AI 中文摘要
情感图像内容分析(AICA)旨在识别和理解视觉内容引发的情感,是通向通用人工智能(AGI)不可或缺的一步。尽管多模态大语言模型(MLLMs)发展迅速,但近期模型发布中对其视觉情商的系统评估仍基本缺失。我们将这一差距归因于传统AICA范式与MLLMs开放式、指令驱动性质之间的结构不匹配。为此,我们引入情感陈述判断(ESJ),一种在保留输入空间表现力的同时将输出约束为判别性判断的陈述验证公式。我们进一步开发了INSETS,通过构建INSETS-462k并支持MVEI(一个涵盖情感极性、情感解释、场景上下文和感知主观性的严格精炼基准)来大规模实例化ESJ。此外,我们构建了EmObserver,一个通过精心设计的多阶段方法在ESJ上优化的面向情感的MLLM。对多种MLLMs在MVEI上的广泛评估揭示了当前人工视觉情商的细粒度见解,而在多个AICA基准上的实验证明了EmObserver的准确性、泛化性和推理忠实性。这些结果确立了ESJ作为一种实用公式、MVEI作为一个综合基准以及EmObserver作为推进面向MLLM的视觉情商的先进基线。代码将在指定链接发布。
英文摘要
Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases. We attribute this gap to a structural mismatch between conventional AICA paradigms and the open-ended, instruction-driven nature of MLLMs, where further analysis reveals four major limitations: omission of plausible responses, limited emotion taxonomies, neglect of contextual factors, and labor-intensive annotation. To overcome these barriers, we introduce Emotion Statement Judgement (ESJ), a statement-verification formulation that preserves the expressiveness of the input space while constraining outputs to discriminative judgements. We further develop INSETS, a labor-efficient pipeline that instantiates ESJ at scale by constructing INSETS-462k and supporting MVEI, a rigorously refined benchmark spanning sentiment polarity, emotion interpretation, scene context, and perception subjectivity. Beyond evaluation, we build EmObserver, an emotion-oriented MLLM optimized on ESJ through an elaborate multi-stage recipe. Extensive evaluation of broad-spectrum MLLMs on MVEI reveals fine-grained insights into current artificial visual emotional intelligence, while experiments on multiple AICA benchmarks demonstrate the accuracy, generalization, and reasoning faithfulness of EmObserver. Collectively, these results establish ESJ as a practical formulation, MVEI as a comprehensive benchmark, and EmObserver as an advanced baseline for advancing MLLM-oriented visual emotional intelligence. Code will be released at: https://github.com/wdqqdw/EmObserver.