arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉语言模型中模态顺序一致性的测试时训练

Test-Time Training for Modality Order Consistency in Vision-Language Models

Aditi Gupta, Yossi Gandelsman

arXiv 2607.20351首次发表:更新:

发表机构

University of Chicago(芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现视觉语言模型对图像和问题呈现顺序敏感,利用此设计测试时训练方法,缩小模态顺序差距,使两种顺序相互一致,定位顺序失败区域,证明该方法可缓解故障并提升性能。

AI 中文摘要

我们发现视觉语言模型对特定语义无关变化敏感:图像和问题呈现的顺序。在三个模型和三个基准测试中,图像优先提示始终优于问题优先提示,揭示了可重复的模态顺序失败。我们利用这一差距设计了一种顺序一致的测试时训练方法。该方法在所有评估设置中大幅缩小了模态顺序差距。令人惊讶的是,它还在更强的图像优先分支上相对于基线产生了一致的改进,使两种顺序相互一致。激活修补将顺序失败定位到网络中间一个狭窄区域,测试时训练方法修复了各层的这种错位。我们的结果表明模态顺序敏感性是视觉语言模型中的电路级故障,并证明简单的非对称测试时适应可以有效缓解它甚至提高性能。

英文摘要

We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a repeatable modality order failure. We use this gap to design an order-consistent test-time training method. Our method substantially closes the modality-order gap across all evaluated settings. Surprisingly, it also yields consistent improvements in the stronger image-first branch over the baseline, hence bootstrapping both orderings toward mutual consistency. Activation patching localizes the ordering failure to a narrow mid-network region where representations diverge sharply between prompt orders. We find that the test-time training method repairs this misalignment across layers. Together, our results identify modality-order sensitivity as a circuit-level failure in VLMs and demonstrate that simple, asymmetric test-time adaptation can effectively mitigate it and even improve performance over the baseline.

Comments16 pages, 7 figures, preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑