arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24730cs.MMcs.HC

学习可靠偏好:带校准融合的错误增强情感偏好优化

Learning to Prefer Reliably: Error-Augmented Emotion Preference Optimization with Calibrated Fusion

Zilong Huang, Junyi Peng, Junjie Li, Kai Li, Wenze Ren, Kong Aik Lee, Man-Wai Mak, Tatsuya Kawahara

中文总结 AI 辅助

针对情感偏好学习的稀疏监督与单一MLLM评判者的偏差问题,提出EAPO框架,构建错误增强数据集并通过边际校准软融合聚合多MLLM评判结果,提升情感偏好预测与鲁棒性。

中文摘要 AI 辅助

情感偏好学习利用候选描述间的成对比较,将多模态大语言模型(MLLMs)与人类对开放式情感描述的判断对齐,并训练捕捉人类情感偏好的奖励模型。然而,传统的成对监督往往较为稀疏,通常仅为每个正面描述提供一个负面描述,因此对情感描述可能出现的各类错误覆盖有限。具体而言,模型可能无法充分接触到语义流畅但情感不一致的描述。除了数据层面的限制,依赖单一MLLM评判者还会引入独特的模型层面问题:其判断在解释细粒度或模糊的多模态情感线索时,可能受到模型特定偏差的影响。为解决这些限制,我们提出错误增强偏好优化(Error-Augmented Preference Optimization, EAPO),这一框架旨在从数据和模型层面提升基于MLLM的情感偏好判断的可靠性。首先,我们从每个偏好描述生成多个受控且感知情感的负面描述,构建错误增强数据集。随后,我们让多个独立的MLLM评判者适应这种更丰富的监督信号,并使用边际校准软融合聚合它们的偏好边际,该方法会在聚合前将异构边际映射到共同尺度。在MER2026-EmoPrefer挑战赛数据集及我们的错误增强数据集上开展的实验表明,EAPO可提升情感偏好预测性能,并增强MLLM评判者在评估与视频多模态情感证据冲突的流畅描述时的鲁棒性。我们的代码可在该https URL获取。

英文摘要

Emotion preference learning uses pairwise comparisons between candidate descriptions to align multimodal large language models (MLLMs) with human judgments of open-ended emotion descriptions and to train reward models that capture human emotional preferences. However, conventional pairwise supervision is often sparse, typically providing only a single negative description for each positive description, and therefore offers limited coverage of the diverse ways in which an emotion description can be incorrect. In particular, models may be insufficiently exposed to semantically fluent but emotionally inconsistent descriptions. Beyond this data-level limitation, relying on a single MLLM judge introduces a distinct model-level concern: its judgments can be affected by model-specific biases when interpreting fine-grained or ambiguous multimodal emotional cues. To address these limitations, we propose Error-Augmented Preference Optimization (EAPO), a framework for improving the reliability of MLLM-based emotion preference judgment at both the data and model levels. First, we construct an error-augmented dataset by generating multiple controlled and emotion-aware negative descriptions from each preferred description. We then adapt multiple independent MLLM judges to this richer supervision and aggregate their preference margins using margin-calibrated soft fusion, which maps heterogeneous margins to a common scale before aggregation. Experiments on the MER2026-EmoPrefer Challenge dataset and our error-augmented dataset demonstrate that EAPO improves emotion preference prediction and enhances the robustness of MLLM judges when evaluating fluent descriptions that conflict with the video's multimodal emotional evidence. Our code is available at https://github.com/slash1028/EAPO-EmoPrefer.

补充信息

↑