发表机构
University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对强化学习在情感推理中忽视主观性与视觉锚定的问题,提出DSPO框架,通过分布对齐奖励和反事实视觉干预,显著提升跨域情感推理稳健性。
AI 中文摘要
强化学习已显著提升了多模态大语言模型(MLLMs)的复杂推理能力。然而,现有的强化学习算法在情感推理任务中遭遇严重失败。这些方法严重依赖确定性硬标签监督和逐点孤立评估,这与人类情感固有的主观性和连续分布特性存在根本性差距。此外,与显式物理对象不同,情感状态在视觉线索中高度隐含。这种抽象性加剧了MLLMs中的视觉幻觉,导致产生看似合理但缺乏依据的情感证据。为解决这些局限,我们提出多样性感知主观策略优化(DSPO),一种联合促进主观情感覆盖和视觉锚定的强化学习框架。首先,我们通过在VAD空间中结合标注情感的词汇先验与图像特定上下文信息,构建上下文锚定的情感分布先验。基于该先验,我们引入分布对齐情感多样性奖励(DEDR),该奖励衡量rollout中每个候选情感的留一法边际贡献。DEDR奖励那些其纳入使预测情感集合更接近上下文锚定先验的候选,从而保留合理的主观解释而不鼓励无约束的分散。我们进一步开发反事实视觉干预门控(CVIG),该机制掩蔽推理过程中被高亮的视觉区域,并利用由此产生的候选级概率变化来降低缺乏视觉证据支持的解释的权重。大量实验表明,DSPO在多个公开基准上达到最先进性能,尤其在跨域性能上,即相比EMO-R3平均跨域准确率提升+10.8%。
英文摘要
Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failure in emotion reasoning tasks. These methods heavily rely on deterministic hard-label supervision and point-wise isolated evaluation, creating a fundamental gap with the inherently subjective and continuously distributed nature of human emotions. Furthermore, unlike explicit physical objects, emotional states are deeply implicit within visual cues. This abstract nature exacerbates visual hallucinations in MLLMs, leading to plausible yet ungrounded emotional evidence. To address these limitations, we propose Diversity-Aware Subjective Policy Optimization (DSPO), a reinforcement learning framework that jointly promotes subjective affective coverage and visual grounding. First, we construct a context-grounded emotional distribution prior in the VAD space by combining the lexical prior of the annotated emotion with image-specific contextual information. Based on this prior, we introduce a Distribution-Aligned Emotional Diversity Reward (DEDR), which measures the leave-one-out marginal contribution of each candidate emotion within a rollout. DEDR rewards candidates whose inclusion brings the predicted affective set closer to the context-grounded prior, thereby preserving plausible subjective interpretations without encouraging unconstrained dispersion. We further develop Counterfactual Visual Intervention Gating (CVIG), which masks the visual region highlighted in the reasoning process and uses the resulting candidate-wise probability changes to reduce the weights of interpretations unsupported by visual evidence. Extensive experiments demonstrate that DSPO achieves state-of-the-art performance across multiple public benchmarks, especially on the cross-domain performance, i.e., improving +10.8\% on average cross-domain accuracy than EMO-R3.
Comments17 pages, 5 figures