arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36782cs.CV

解码情感细微差别:通过层级情感推理与对比判别剪枝增强多模态大语言模型

Decoding Affective Nuances: Enhancing MLLMs via Hierarchical Emotion Reasoning and Contrastive Discriminative Pruning

Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu, Zhendong Mao

首次发表
浏览论文内容

中文总结 AI 辅助

提出DAN框架,通过层级情感推理链和对比判别视觉剪枝,无需训练即可增强多模态大模型对细微情感的判别能力,在WebEmo25上提升10.47%。

中文摘要 AI 辅助

尽管多模态大语言模型(MLLMs)在客观理解任务中展现出卓越的能力,但其在情感推理方面的表现仍远未达到人类水平。我们将此归因于一个核心能力差距:MLLMs难以基于细粒度视觉证据可靠地区分语义相近的情感,这可以分解为两个局限:1)归因不足。传统MLLMs的全局推理范式严重稀释了细粒度情感线索,其中微妙的情感状态通常被隐式编码,从而在复杂场景中产生情感误判。2)判别不足。现有方法只能识别与情感大致相关的区域,无法区分语义相似情感之间的判别性区域,导致情感判断模糊。为克服这些局限,我们提出了一种无需训练的推理时优化框架,名为解码情感细微差别(DAN)。具体而言,我们提出层级情感推理链(HERC),通过协调细粒度场景/对象级线索并执行软门控推理来增强归因不足。此外,为了区分语义相近的情感,我们设计了对比判别视觉剪枝(CDVP),通过计算相似情感注意力分布之间的绝对差异来隔离判别性视觉标记,以推理最终情感类别。在多个基准上的性能表明,DAN显著提高了对情感细微差别的判别能力,且无需消耗额外训练资源,尤其在WebEmo25数据集(包含25个细粒度情感类别)上,使用Qwen3-VL-8B-Instruct实现了+10.47%的提升。

英文摘要

While multimodal large language models (MLLMs) have demonstrated exceptional capabilities in objective understanding tasks, their performance in affective reasoning still falls significantly short of human standards. We attribute it to a central capability gap: MLLMs are difficult to reliably distinguish semantically proximal emotions based on fine-grained visual evidence, which could be decoupled as two limitations: 1) Insufficient Attribution. The global reasoning paradigm of conventional MLLMs severely dilutes fine-grained emotion cues, where subtle emotional states are usually implicitly encoded, thereby generating emotional misjudgments in complex scenarios. 2) Insufficient Discrimination. Existing methods could only identify regions generally associated with emotions, which fails to distinguish discriminative regions between semantically similar emotions, leading to ambiguous emotion judgements. To overcome these limitations, we present a training-free inference-time optimization framework, named Decoding Affective Nuances (DAN). Specifically, we propose a Hierarchical Emotional Reasoning Chain (HERC) that enhances the insufficient attribution by harmonizing fine-grained scene/object-level cues and performing a soft-gated reasoning. Furthermore, to discriminate between semantically proximal emotions, we design a Contrastive Discriminative Visual Pruning (CDVP), which isolates discriminative visual tokens to reason the final emotion category by computing the absolute discrepancy between the attention distributions of similar emotions. Performances on several benchmarks demonstrate that DAN significantly improves discrimination for affective nuances without consuming additional training resources, especially achieving +10.47% improvements with Qwen3-VL-8B-Instruct on WebEmo25 dataset that contains 25 fine-grained emotion categories.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Harbin Institute of Technology, Weihai(哈尔滨工业大学(威海))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑