arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁在谁的左边?追踪相对位置推理中的空间证据与角色绑定

Who Is Left of Whom? Tracing Spatial Evidence and Role Binding in Relative-Position Reasoning

Yingjin Song, Denis Paperno, Albert Gatt

arXiv 2609.35486首次发表:更新:

发表机构

Utrecht University(乌得勒支大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过激活修补和定向干预,揭示视觉语言模型在相对位置推理中分阶段处理空间证据与角色绑定,并提出可泛化的干预方向以提高推理一致性。

AI 中文摘要

高实例级准确率可能掩盖空间推理中的不一致性,当对象交换位置或查询中角色反转时,这种不一致性尤为明显。支撑相对位置推理的内部表征仍鲜为人知。我们研究了该过程的两个互补组成部分:追踪输入中的对象位置以及表征其在查询中的角色。在三个视觉语言模型(VLM)及其语言模型骨干上,分别使用视觉或文本输入,激活修补揭示了一个分阶段进展:从早期层的源表征,经过中间层的查询对象表征,到后期层的答案状态。定向干预进一步确立了沿此进展的因果联系:操纵源侧表征会转移查询对象提及处的空间信息,并最终改变关系预测。除对象位置信息外,我们还识别出一个与比较中两个对象角色相关的稳定查询侧方向。在合成场景上估计的方向可泛化至自然图像基准,在多数设置中无需重新训练即可提高准确率及两种形式的配对一致性。我们的发现揭示了视觉与文本设置中关系推理的互补组成部分,并展示了定向干预如何提升模型行为的一致性。

英文摘要

High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate two complementary components of this process: tracking object locations in the input and representing their query roles. Across three VLMs with visual or textual inputs and their language-model backbones, activation patching reveals a staged progression from early-layer source representations through intermediate-layer query-object representations to late-layer answer states. Targeted interventions further establish causal links along this progression: manipulating source-side representations shifts location information at query-object mentions and ultimately alters relation predictions. Beyond object-location information, we also identify a stable query-side direction associated with the roles of the two objects in the comparison. Steering along directions estimated on synthetic scenes generalizes to natural-image benchmarks, improving accuracy and both forms of paired consistency in most settings without retraining. Our findings reveal complementary components of relational reasoning across visual and textual settings and show how targeted interventions can improve the consistency of models' behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑