arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ArtECulture:评测多模态大语言模型中的文化条件视觉情感理解

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

Xiaolin Chen, Xuemeng Song, Wenhao Shi, Xianjing Han, Mong-Li Lee, Wynne Hsu

arXiv 2608.03358首次发表:更新:

AI 中文总结

针对现有视觉情感理解方法忽略文化差异的问题,提出ArtECulture基准与检索增强型文化条件情感理解框架,评估多模态大语言模型在该任务上的表现并验证框架的提升效果。

AI 中文摘要

现有视觉情感理解方法通常忽略情感感知中的文化差异。我们提出文化条件视觉情感理解任务,即预测给定图像对应的特定文化情感感知并解释其潜在原理。尽管已有相关基准,但它们存在个体标注不一致、难以生成多数支持的文化层面情感标签、文化覆盖不均衡的问题。因此,我们推出ArtECulture基准,包含6792件艺术品,涵盖英语、汉语、阿拉伯语文化的特定文化情感标签与解释,西方与非西方内容均衡。对16个开源与闭源多模态大语言模型(MLLMs)的零样本评估显示,该任务仍具挑战性,最佳模型准确率低于50%。为解决此局限,我们提出检索增强型文化条件情感理解框架,利用基于概念的文化情感知识库,在无需额外训练的情况下将显式文化知识注入MLLMs,该框架可提升文化对齐情感预测与基于依据的解释生成。我们的基准与代码将公开发布。

英文摘要

Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task that predicts the culture-specific emotional perception of a given image and explains the underlying rationale. Although related benchmarks exist, they are limited by inconsistent individual annotations, which hinder the derivation of majority-supported culture-level emotion labels, and imbalanced cultural coverage. Thus, we present ArtECulture, a benchmark containing 6,792 artworks with culture-specific emotion labels and explanations across English, Chinese, and Arabic cultures, with balanced Western and non-Western content. Evaluations of 16 open- and closed-source Multimodal Large Language Models (MLLMs) under a zero-shot setting reveal that the task remains challenging, with the best model achieving below 50\% accuracy. To address this limitation, we introduce a retrieval-augmented culture-conditioned emotion understanding framework, which leverages a concept-based cultural emotion knowledge base to inject explicit cultural knowledge into MLLMs without additional training. The framework improves both culturally aligned emotion prediction and grounded explanation generation. Our benchmark and code will be publicly released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑