基于视觉语言模型的不确定性感知艺术史年代测定
Uncertainty-Aware Art-Historical Dating with Vision-Language Models
浏览论文内容
中文总结 AI 辅助
该研究针对艺术品年代测定中的时间纠缠问题,将其转化为不确定性感知回归任务,通过在Wikidata艺术品语料库上实验发现视觉语言模型(VLMs)在该任务上优于纯视觉自监督基线,但存在偏差。
中文摘要 AI 辅助
博物馆和档案数据集并不能反映历史艺术创作,而是体现了收藏、保存、编目和数字化的偶然历史,这对预训练图像表征的解读有直接影响:这些表征看似编码了历史时间,实则编码了物体成为可见数据的制度条件,我们将此现象称为时间纠缠。我们将艺术品年代测定表述为针对冻结图像嵌入的不确定性感知回归任务,以此研究该现象。我们在经过时间控制的Wikidata艺术品语料库上评估了多种预训练视觉模型,结果显示这些模型包含可用的时间信息,其中视觉语言模型(VLMs)的表现优于纯视觉自监督基线模型。不过,定性分析表明,这种时间知识受多种偏差影响。
英文摘要
Museum and archival datasets do not mirror historical artistic production, but materialize the contingent histories of collecting, preservation, cataloging, and digitization. This has direct consequences for interpreting pretrained image representations: they may appear to encode historical time while actually encoding the institutional conditions under which objects become visible as data. We describe this phenomenon as temporal entanglement and investigate it by formulating artwork dating as an uncertainty-aware regression task over frozen image embeddings. We evaluate several pretrained vision models on a temporally controlled Wikidata corpus of artworks. Our results show that these models contain usable temporal information, with Vision-Language Models (VLMs) outperforming purely visual self-supervised baselines. However, a qualitative analysis indicates that this temporal knowledge is shaped by various biases.
发表机构
- Marburg University(马尔堡大学)
机构由 AI 辅助整理,请以论文原文为准。