arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

先验方向:为什么GUI grounding会被锁定在过去

Prior Directions: Separating Geometric Identification from Behavioral Efficacy in Visual Revision

Weile Gong, Zijian Lu, Mingcai Chen, Yiping Zuo, Xin He, Weibei Fan

arXiv 2607.26913首次发表:更新:

发表机构

Nanjing University of Posts and Telecommunications(南京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探究视觉-语言模型的GUI grounding锁定问题,提出先验方向概念,发现其是锁定的关键,移除该方向分量可恢复视觉grounding。

AI 中文摘要

视觉-语言模型常利用早期视觉状态的描述对当前场景做出决策。当场景发生变化时,过时的语言会将原本正确的视觉判断导向过时的答案。我们在受控的grounding场景中研究这种被称为视觉锁定的失败,其中仅口头表述的先验发生变化。在所有模型中,更强的锁定伴随着最终答案前模型表示的更小变化。这种反转表明,锁定不取决于该表示移动的距离,而取决于该移动的组织方式。在更难纠正的模型中,先验诱导的变化集中在跨示例重复出现的紧凑方向集合上,我们将这些重复出现的轴称为先验方向(Prior Directions)。它们在未见过的示例上重复出现,而对四个模型的描述性比较表明,更高的集中度与更强的锁定相关。受控干预显示,移除与先验方向对齐的分量可恢复视觉grounding,而移除同等大小的正交分量几乎没有效果。因此,先验控制产生于先验诱导的变化在用于生成答案的表示中形成连贯且可重复使用的模式,这一解释说明了为何同一先验在一个模型中仍可修正,而在另一个模型中却占主导地位。

英文摘要

Vision-language models can remain anchored to stale verbal priors even when current visual evidence supports a different answer. We use this controlled revision failure to ask whether the subspace that best captures recurrent prior-induced representation change is also uniquely more behaviorally effective. Across four models, paired prior and reference states define Prior Directions, a compact low-rank geometry that, at rank 32, spans only 0.78% of the state dimensions yet captures 49.7% to 64.0% of held-out displacement energy. Matched correspondence improves held-out capture at every tested rank in all four models; across 20 permutation seeds, this advantage remains positive in all 15 available combinations of model and rank. Yet the geometric advantage does not translate into a comparable intervention advantage: matched and Pairing-Permuted edits restore 60/66 and 59/66 clean prior-induced failures, respectively. We trace this discrepancy to the geometry of the realized edits. Pairing-Permuted edits remain strongly aligned with the matched Prior Directions; their aligned component restores 59/66 cases, whereas an equally norm-matched orthogonal residual restores only 20/66. These results show that the construction that best identifies recurrent representation change need not be uniquely more behaviorally effective, distinguishing global subspace identification from the geometry actually used by an intervention.

CommentsCode: https://github.com/phare111/prior-directions

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑