超越语义的艺术:用于多关系表示的层状信息对比学习
Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations
- University of Zurich(苏黎世大学)
- Max Planck Institute Bibliotheca Hertziana(马克斯·普朗克赫兹iana图书馆研究所)
- Sapienza University of Rome(罗马第一大学)
- Amazon(亚马逊)
- University of Amsterdam(阿姆斯特丹大学)
- The University of Osaka(大阪大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对视觉语言模型丢失艺术作品多关系结构的问题,引入受层理论启发的CANVAS框架,通过多嵌入和对比损失学习关系感知多模态表示,在新基准测试中表现优于基线,证明多关系对齐在艺术理解中兼具理论与实践价值。
AI中文摘要:
理解一幅画绝非单一行为。艺术史学家会通过风格、图像学或历史背景等概念来分析同一作品,这些维度不可互换,且各自在视觉与文本间承载着不同语义关系。像CLIP这样的视觉语言模型(VLMs)学习单一共享嵌入空间,将这种丰富性简化为单一均匀对齐,从而丢失了定义艺术史推理的多关系结构。我们引入了CANVAS(基于层的视觉语言对齐对比艺术感知网络),这是一个受层理论启发用于学习关系感知多模态表示的框架。每个艺术品根据关系类型(即上下文)投影到多个嵌入中,一种新颖的对比损失在训练期间编码上下文信息,推理时不依赖外部数据。我们在三个新引入的用于多关系艺术理解的艺术品基准上进行评估:源自WikiArt和维基百科的WikiArt +、来自赫茨iana图书馆藏品的HertzianaDP以及从SemArt数据集提炼的SemArt +。在多模态检索和艺术理解方面,CANVAS优于基线,支持了多关系对齐不仅在理论上有动机而且在实践中也至关重要的观点。
英文摘要:
Understanding a painting is never a single act. Art historians may analyze the same work through concepts of style, iconography, or historical context, dimensions that are not interchangeable, and each carries distinct semantic relationships between the visual and the textual. Vision-Language Models (VLMs) like CLIP, which learn a single shared embedding space, collapse this richness into a single homogeneous alignment, thereby losing the multi-relational structure that defines art-historical reasoning. We introduce CANVAS (Contrastive Art-aware Network for Vision-Language Alignment with Sheaves), a framework for learning relation-aware multimodal representations inspired by sheaf theory. Each artwork is projected into multiple embeddings conditioned on the type of relation (i.e., the context), and a novel contrastive loss encodes contextual information during training, with no dependency on external data at inference. We evaluate on three newly introduced benchmarks of artworks for multi-relational art understanding: WikiArt+, derived from WikiArt and Wikipedia, HertzianaDP, from the Bibliotheca Hertziana collection, and SemArt+, refined from the SemArt dataset. In multimodal retrieval and art understanding, CANVAS outperforms the baselines, supporting the view that multi-relational alignment is not just theoretically motivated but also practically essential.