发表机构
University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences(中国科学院大学; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多模态情感理解中EmoPrefer评估指标,通过内容盲探针审计发现仅用描述长度和生成器标识的逻辑回归性能与复杂模型相当,揭示当前得分可不验证描述与视频达成,建议未来评估采用源平衡配对等方法。
AI 中文摘要
对模型生成的情感描述的偏好正成为多模态情感理解的标准评估指标,如EmoPrefer中的MER2026 MER-Prefer赛道。此类基准假设预测首选描述需要对视频有基于内容的跨模态理解。我们使用内容盲探针对EmoPrefer进行了系统的捷径审计。仅使用描述长度和生成器标识的简单逻辑回归,在不处理文本、视频或音频的情况下,性能与LoRA微调的7B文本和视听评判相当。生成器标识可从描述文本中以99.5%的准确率恢复,每个候选对对比两个不同的生成器,人类偏好标签在66%的评估对上与每个生成器的独占胜率先验一致。当人类标签与该先验冲突时,训练有素的评判在63%至80%的对上仍遵循风格先验。在消除冗长偏差的长度匹配子集中,测试的媒体配置没有统计学上的显著改善,而一种受ODIN启发的诊断方法,将风格捷径解耦后,其内容得分接近随机。这些结果并不意味着人类偏好本质上是风格性的,也不意味着描述中不包含情感信息。相反,它们表明当前得分可以在不将任何描述与视频进行验证的情况下达到。我们建议在未来的跨生成器评估中采用源平衡配对、严格的长度控制、反刻板切片报告和多注释者共识。代码可在这个https URL获取。
英文摘要
Preference over model-generated emotion descriptions is emerging as a standard evaluation metric for multimodal emotion understanding, exemplified by the MER2026 MER-Prefer track on EmoPrefer. Such benchmarks assume that predicting the preferred description requires grounded cross-modal understanding of the video. We conduct a systematic shortcut audit of EmoPrefer using content-blind probes. A simple logistic regression using only description length and generator identity, without processing the text, video, or audio, performs comparably to LoRA-finetuned 7B text and audio-visual judges (65.8 versus 66.8 WAF on EmoPrefer-V2). Generator identity is recoverable from description text with 99.5 percent accuracy, every candidate pair contrasts two distinct generators, and the human preference labels agree with a fold-exclusive per-generator win-rate prior on 66 percent of the evaluated pairs. When the human label conflicts with this prior, trained judges still follow the style prior on 63 to 80 percent of the pairs. On a length-matched subset that neutralizes verbosity bias, the tested media configurations yield no statistically significant improvement, while an ODIN-inspired diagnostic that decouples the style shortcut leaves its content head near chance. These results do not imply that human preferences are inherently stylistic or that the descriptions contain no emotional information. Instead, they show that the current scores can be reached without verifying either description against the video. We recommend source-balanced pairing, strict length control, counter-stereotypical sliced reporting, and multi-annotator consensus for future cross-generator evaluations. Code is available at https://github.com/jiabingyang01/EmoPrefer-Audit.