arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26168cs.HCcs.CV

TRACE:抽象概念评估的透明检索

TRACE: Transparent Retrieval for Abstract Concept Evaluation

Joseph Bingham

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出透明检索基线,使用SIFT匹配和信号质量指数,在七巧板指称任务中达到与最强VLM相当的性能,表明学习视觉表示并非瓶颈,且VLM间能力差异显著。

中文摘要 AI 辅助

近期研究报道,视觉-语言模型(VLM)在重复指称游戏中难以建立和维持稳定的指称。我们并非追问哪个VLM表现最佳,而是提出一个更基本的问题:对于这一任务,是否真的需要大型预训练VLM?在将单个导演话语落地到十二个七巧板轮廓之一的任务中,我们比较了六种现成的VLM与一个透明基线,该基线不使用任何学习到的视觉表示:经典的SIFT关键点匹配和基于检索图像的信号质量指数。在相同试验中,透明基线匹配了最强的VLM(SigLIP-large),并显著优于其他五种,包括所有CLIP和OpenCLIP变体。该基线额外检索外部图像,因此这不是信息匹配的比较;它表明学习到的视觉表示并非此任务的瓶颈:给定检索图像,形状匹配的经典相似度就足够了。在此过程中,我们发现抽象落地能力在VLM之间差异很大(top-1准确率15%-39%;随机水平8.33%,人类约77%-80%),因此弱点特定于模型,而非对比预训练的内在缺陷;在包含1,013个形状的KiloGram基准上,该模式对CLIP具有普遍性,每个形状的难度与人类对形状的可命名性相关。该流程是一种经典的、可检查的替代方案,而非学习型方案。最后,我们概述了如何通过显式、可检查的听者侧契约状态表示,将这种方法扩展到交互式多轮指称,这留待未来工作。代码见补充材料。

英文摘要

Recent work reports that vision--language models (VLMs) struggle to establish and maintain stable reference in repeated reference games. Rather than ask which VLM does best, we ask a more basic question: do you need a large pretrained VLM for this at all? On grounding a single director utterance to one of twelve tangram silhouettes, we compare six off-the-shelf VLMs against a transparent baseline that uses \emph{no learned visual representation}: classical SIFT keypoint matching and a signal-quality index over retrieved images. On identical trials, the transparent baseline matches the strongest VLM (SigLIP-large) and significantly outperforms the other five, including every CLIP and OpenCLIP variant. The baseline additionally retrieves external images, so this is not a matched-information comparison; what it shows is that a learned \emph{visual} representation is not the bottleneck for this task: given retrieved images, a shape-appropriate classical similarity suffices. Along the way we find that abstract-grounding ability varies widely across VLMs (15--39\% top-1; chance 8.33\%, humans $\approx$77--80\%), so the weakness is model-specific rather than intrinsic to contrastive pretraining; on the 1{,}013-shape KiloGram benchmark the pattern generalizes for CLIP, with per-shape difficulty tracking human shape-nameability. The pipeline is a classical, inspectable alternative rather than a learned one. We close by sketching how an explicit, inspectable representation of listener-side pact state could carry this approach into interactive multi-turn reference, which we leave to future work. Code available in supplementary material.

发表机构

  • Technion University(以色列理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑