AI 中文总结
研究探讨用AAT词汇表的根面类型和层次深度解释CLIP在历史照片集检索的成败及微调作用,通过三个标注集研究视觉连贯性等指标,发现两指标可区分失败模式,根面类型有区分作用,微调对层次浅的术语更有效。
AI 中文摘要
GLAM机构(美术馆、图书馆、档案馆和博物馆)使用如艺术与建筑词库(AAT)等控制词汇表来组织图像访问。在基于内容的图像检索中,CLIP等视觉语言模型的性能各异,与概念抽象和概念频率等有关,但此前尚无工作从AAT等词汇表的结构属性解释这种差异。AAT将概念分组并分层排列。本文探讨根面类型和层次深度这两个结构属性能否解释CLIP检索的成败及微调的作用。通过三个用AAT标注的历史照片集,研究视觉连贯性、文本-图像对齐及标准检索指标,发现视觉连贯性和文本-图像对齐几乎不相关且能区分不同失败模式,根面类型能显著区分视觉连贯性不同的类别,微调总体上能改善检索,对层次中较浅的术语效果更好。
英文摘要
GLAM institutions (Galleries, Libraries, Archives, and Museums) organise image access using controlled vocabularies such as the Art and Architecture Thesaurus (AAT). For content-based image retrieval in these settings, vision-language models like CLIP are increasingly used, but their performance varies. This variation is known to relate to measures like concept abstraction and concept frequency. However, no prior work explains this variation in terms of the structural properties of vocabularies like the AAT that GLAM professionals already use. The AAT groups concepts into broad facets (Objects, Activities, Agents, etc.) and arranges terms hierarchically within them. In this paper, we ask whether two structural properties (root facet type and hierarchy depth) explain where CLIP retrieval succeeds and fails, and where fine-tuning helps. Across three historical photographic collections annotated with AAT terms, we examine visual coherence (whether a term's photographs cluster in CLIP's embedding space), text-image alignment (whether its label is near that cluster), and standard retrieval measures, which conflate the two. We find that visual coherence and text-image alignment are nearly uncorrelated across terms and jointly separate distinct failure modes. Terms whose photographs cluster tightly but whose label is distant from the cluster retrieve poorly in every collection, in two of three collections even worse than terms that fail on both metrics. We also show that while retrieval metrics do not correlate significantly with either structural property, root facet type does significantly separate categories with varying visual coherence. Finally, we find that fine-tuning improves retrieval overall, but its gains favour shallower terms in the hierarchy, where text-image alignment improves most, beyond what concept frequency explains.