arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向艺术品与照片的组合分析学习视觉表征

Learning visual representations for compositional analysis of artworks and photographs

Fatemeh Behrad, Tinne Tuytelaars, Johan Wagemans

arXiv 2608.06142首次发表:更新:

发表机构

KU Leuven University(KU Leuven大学(荷语鲁汶大学))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对比了受人类启发的构图分析方法与微调基础模型两种范式,在三类任务上评估性能,发现前者冻结编码器时具竞争力且可解释,后者微调后性能更优但可解释性与泛化性下降。

AI 中文摘要

构图是对视觉元素的刻意安排,是艺术品中意义、情感与美学品质传递的核心,但它仍是视觉理解中最未被形式化的维度之一。现有研究指出,学习有意义的构图表征存在持续的差距,将其归因于语义偏差,并提出受人类启发的方法可能是关键。我们比较了两种并行的构图分析范式:一种是基于感知分组的受人类启发的方法,另一种是利用近期大规模构图数据集实现的微调基础模型。受人类启发的方法采用以对象为中心的模型进行区域级分解,并使用图注意力网络捕捉元素间的空间关系。两种范式均在构图分数/类别预测、构图图像检索和视觉显著性检测任务上进行评估。在冻结编码器的情况下,受人类启发的方法实现了有竞争力的性能,同时保持了可解释性;当有足够数据进行微调时,大型自监督模型的性能显著更优,但代价是可解释性和跨域泛化能力的下降。

英文摘要

Composition, the deliberate arrangement of visual elements, is central to how meaning, emotion, and aesthetic quality are conveyed in artwork, yet it remains among the least formalized dimensions of visual understanding. Prior work highlights a persistent gap in learning meaningful compositional representations, attributing it to semantic bias and suggesting that human-inspired approaches may be key. We compare two parallel paradigms for composition analysis: a human-inspired method grounded in perceptual grouping, and fine-tuned foundation models enabled by recent large-scale compositional datasets. The human-inspired approach uses object-centric models for region-level decomposition and a graph attention network to capture spatial relationships between elements. Both paradigms are evaluated on composition score/category prediction, compositional image retrieval, and visual saliency detection. With frozen encoders, the human-inspired method achieves competitive performance while remaining interpretable. When sufficient data enables fine-tuning, large self-supervised models outperform significantly, but at the cost of interpretability and cross-domain generalization. Code and pre-trained models are available on GitHub.

CommentsECCV workshops 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑