arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10706cs.CVcs.MM

MMArt:用于视觉艺术理解的多视角多模态数据集

MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding

Shuai Wang, Wangyuan Ding, Yixian Shen, Jia-Hong Huang, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg, Marcel Worring

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有艺术数据集单视角局限,推出含74234幅WikiArt画作的多视角多模态数据集MMArt,通过分析各视角特性证实多视角设计的必要性,为视觉艺术理解提供支撑。

中文摘要 AI 辅助

近期的视觉-语言模型展现出出色的通用视觉理解能力,但其艺术解读仍较为浅显:它们仅能描述表面内容,难以进行形式分析、基于语境的历史解读或情感刻画。我们认为这不仅是模型的问题,也是数据集的局限。现有艺术数据集均为单视角资源,尚无数据集能针对同一艺术作品同时提供叙事、形式、情感和历史视角的标注。我们推出MMArt,这是一个包含74234幅WikiArt画作的大规模数据集,每幅画作都带有四个独立标注的视角,以及由专用视觉-语言模型或人工标注生成并经互补质量评估验证的统一对齐标题。两项互补分析证实各视角编码了真正不同的信息:生成分析显示,形式分析描述最能保留构图风格,历史描述在重构图像中带有强烈的情感信号;判别式检索分析揭示了任务不对称性:叙事描述驱动检索(R@1=44.0%),而形式描述虽重构能力最强,在检索规模上几乎不具判别性(R@1=7.8%);留一法分析进一步确认,在两项任务中历史描述是最不可替代的视角。两项分析共同表明,单一视角无法满足所有任务需求,直接推动了MMArt的多视角设计。该数据集、代码及更多信息可在指定URL获取。

英文摘要

Recent vision-language models demonstrate impressive general visual understanding, yet their art interpretation remains shallow: they describe surface content but struggle with formal analysis, grounded historical interpretation, or affective characterization. We argue this is not only a model but also a dataset limitation. Existing art datasets are single perspective resources, where no dataset provides narrative, formal, emotional, and historical perspectives simultaneously for the same artworks. We introduce MMArt, a large-scale dataset of 74,234 WikiArt paintings, each annotated with four independently annotated perspectives plus a harmonized unified caption, produced by specialized vision-language models or human annotation and validated through complementary quality evaluations. Two complementarity analyses establish that perspectives encode genuinely distinct information. A generative analysis shows that formal analysis descriptions best preserve compositional style, and historical descriptions carry strong affective signal in reconstructed images. A discriminative retrieval analysis reveals task-asymmetry: narrative descriptions drive retrieval (R@1 = 44.0%), while formal descriptions, strongest for reconstruction, are nearly nondiscriminative at retrieval scale (R@1 = 7.8%). Leave-one-out analysis further confirms that historical descriptions are the least replaceable perspective across both tasks. Together, the two analyses establish that no single perspective suffices for all tasks, directly motivating MMArt multi-perspective design. The dataset, code, and additional information are available at https://shuaiwang97.github.io/MMArt/.

发表机构

  • University of Amsterdam(阿姆斯特丹大学)
  • Amazon AGI(亚马逊AGI)
  • College of Business and Economics, University of Johannesburg(约翰内斯堡大学经济与商学院)

机构由 AI 辅助整理,请以论文原文为准。

↑