arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21012cs.CVcs.LG

面向壁画残片风格分类的片段感知视觉Transformer

Fragment-Aware Vision Transformers for Fresco-Fragment Style Classification

Sara Miketek, Biagio Barchielli, Nadeem Iqbal Kajla, Sinem Aslan

首次发表
浏览论文内容

中文总结 AI 辅助

提出片段感知ViT框架,通过前景掩码、修复正则化和分布级对比学习提升壁画残片风格分类,集成模型在CLEOPATRA上准确率达0.656。

中文摘要 AI 辅助

艺术风格分类通常在完整艺术品上进行研究,此时模型可以利用全局构图、空间组织和图像学结构。然而,在考古场景中,艺术品往往仅以碎片形式存留,迫使识别依赖于不完整、不规则且上下文受限的视觉证据。我们使用一种渐进式基于Transformer的框架研究壁画残片风格分类。从ViT-B/16基线出发,我们引入前景引导掩码以抑制仅含背景的令牌,基于修复的几何正则化以将不规则碎片支撑与ViT补丁网格对齐,以及一个监督对比目标,该目标通过Kullback-Leibler相似性作用于预测分布,并持续改进每个分支。我们通过一个刻意简单的可学习logit集成来组合各分支。在CLEOPATRA和POMPAAF上的实验表明,片段感知建模优于标准ViT基线,集成在CLEOPATRA上将准确率从0.604提升至0.656,宏F1从0.596提升至0.648,并在POMPAAF的六种碎片化设置中的四种上优于最佳单分支。我们还评估了一种更复杂的图融合变体,发现它在POMPAAF上与简单集成相当,而在CLEOPATRA上仅带来微小的、特定于数据集的增益,这不足以证明其增加的复杂性。除了这些经验性增益,我们的贡献有两方面:一个在分布层面持续锐化单分支识别的对比目标,以及一个可解释性分析,该分析验证了模型利用真实的绘画证据,同时量化了基于修复的分支将其部分归因归因于合成环绕区域。

英文摘要

Artistic style classification is usually studied on complete artworks, where models can exploit global composition, spatial organisation, and iconographic structure. In archaeological settings, however, artworks often survive only as fragmented remains, forcing recognition from incomplete, irregular, and context-limited visual evidence. We study fresco-fragment style classification using a progressive transformer-based framework. Starting from a ViT-B/16 baseline, we introduce foreground-guided masking to suppress background-only tokens, inpainting-based geometric regularisation to align irregular fragment supports with the ViT patch grid, and a supervised contrastive objective that operates on predictive distributions through a Kullback-Leibler similarity and consistently improves every branch. We combine the branches with a deliberately simple learnable logit ensemble. Experiments on CLEOPATRA and POMPAAF show that fragment-aware modelling improves over the standard ViT baseline, with the ensemble increasing accuracy from 0.604 to 0.656 and macro-F1 from 0.596 to 0.648 on CLEOPATRA, and outperforming the best single branch in four of six fragmentation settings on POMPAAF. We additionally evaluate a more complex graph-fusion variant and find that it matches the simple ensemble on POMPAAF while offering only a small, dataset-specific gain on CLEOPATRA, which does not justify its added complexity. Beyond these empirical gains, our contribution is twofold: a distribution-level contrastive objective that consistently sharpens single-branch recognition, and an interpretability analysis that verifies the models exploit genuine painted evidence, while quantifying that the inpainting-based branch draws part of its attribution from the synthesised surround.

发表机构

  • Ca’ Foscari University of Venice(威尼斯大学)
  • Dundalk Institute of Technology(邓多克理工学院)
  • University of Milan(米兰大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑