发表机构
University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现视觉语言模型对图像和问题呈现顺序敏感,利用此设计测试时训练方法,缩小模态顺序差距,使两种顺序相互一致,定位顺序失败区域,证明该方法可缓解故障并提升性能。
AI 中文摘要
我们发现视觉语言模型对特定语义无关变化敏感:图像和问题呈现的顺序。在三个模型和三个基准测试中,图像优先提示始终优于问题优先提示,揭示了可重复的模态顺序失败。我们利用这一差距设计了一种顺序一致的测试时训练方法。该方法在所有评估设置中大幅缩小了模态顺序差距。令人惊讶的是,它还在更强的图像优先分支上相对于基线产生了一致的改进,使两种顺序相互一致。激活修补将顺序失败定位到网络中间一个狭窄区域,测试时训练方法修复了各层的这种错位。我们的结果表明模态顺序敏感性是视觉语言模型中的电路级故障,并证明简单的非对称测试时适应可以有效缓解它甚至提高性能。
英文摘要
We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a repeatable modality order failure. We use this gap to design an order-consistent test-time training method. Our method substantially closes the modality-order gap across all evaluated settings. Surprisingly, it also yields consistent improvements in the stronger image-first branch over the baseline, hence bootstrapping both orderings toward mutual consistency. Activation patching localizes the ordering failure to a narrow mid-network region where representations diverge sharply between prompt orders. We find that the test-time training method repairs this misalignment across layers. Together, our results identify modality-order sensitivity as a circuit-level failure in VLMs and demonstrate that simple, asymmetric test-time adaptation can effectively mitigate it and even improve performance over the baseline.
Comments16 pages, 7 figures, preprint