发表机构
ELLIS Institute Finland; Aalto University; University of Oulu; Nanyang Technological University(ELLIS芬兰研究所; 阿尔托大学; 奥卢大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉语言模型在定向空间关系上得分接近随机的问题,提出反对称位移读出方法,利用冻结特征恢复方向信息,无需额外训练,显著优于现有分数并具计算优势。
AI 中文摘要
CLIP类视觉语言模型仍然是多模态系统的基石,然而它们在定向空间关系上的得分仍接近随机水平,例如判断一个物体是否在另一个物体的左侧。我们将这种失败称为读出盲区,并从理论和实证上分析为何部署的分数会遗漏方向:当评分规则对称地对待主语和宾语时,无论编码器如何训练,方向都会被抵消。在此分析的指导下,我们引入了反对称位移读出(ADR),它将冻结特征中的标题词与图像块对齐,并通过匹配对象质心之间的有符号位移来对每个关系进行评分。值得注意的是,ADR无需额外训练或学习参数即可成功,从而证明方向信息仍保留在冻结编码器中。然而,文本和世界先验可能会夸大准确性,因此我们进一步引入了先验去膨胀,它通过将图像-文本配对与将每个项目与无关图像配对的空模型进行比较,来衡量配对的接地增益作为基础增益。跨编码器家族的广泛实验表明,ADR显著优于部署的分数,后者在大多数方向平衡的数据集上即使对于微调编码器仍接近随机水平。与更复杂的读出方法相比,ADR优于所评估的MLLM似然读出方法,并且以一小部分计算量即可与它们的聊天推理相媲美。这些结果支持我们的主张,即通过适当的读出可以从冻结特征中恢复方向信息。我们的实现和评估工具包将公开发布。
英文摘要
CLIP-like vision-language models remain a cornerstone of multimodal systems, yet their scores stay near chance on directed spatial relations, such as whether one object is left of another. We call this failure readout blindness and analyze, theoretically and empirically, why deployed scores miss the direction: when scoring rules treat the subject and object symmetrically, direction cancels regardless of encoder training. Guided by this analysis, we introduce Antisymmetric Displacement Readout (ADR), which aligns caption words with image patches in the frozen features and scores each relation by the signed displacement between matched object centroids. Notably, ADR succeeds without additional training or learned parameters, thereby demonstrating that directional information remains in the frozen encoder. However, text and world priors can inflate accuracy, so we further introduce prior deflation, which measures the benefit of the image-text pairing as the grounded gain over a null that pairs each item with an unrelated image. Extensive experiments across encoder families show that ADR substantially improves over deployed scores, which remain near chance on most direction-balanced sets even for fine-tuned encoders. Compared with more complex readouts, ADR outperforms the evaluated MLLM likelihood readouts and is competitive with their chat inference at a small fraction of the computation. These results support our claim that directional information can be recovered from frozen features by an appropriate readout. Our implementation and evaluation kit will be publicly available.