arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

问两次,看两次:提示回声解决视觉语言模型中的问题优先悖论

Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models

Rakshanda Hassan Abhinandan, John Galeotti, Deva Ramanan, Gautam Rajendrakumar Gare

arXiv 2607.15565首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言模型中问题优先提示的悖论,发现问题前置虽能引导感知但后续答案生成阶段难关注到问题。提出问题回声方法,即在图像两侧重复问题,解决了悖论,在多基准测试中表现出色且无需训练等。

AI 中文摘要

在视觉语言模型(VLM)提示中,问题应置于图像之前还是之后?直觉上应在之前,因为知道问题能引导模型关注方向。然而在视觉问答基准测试中,问题优先提示始终不如前沿VLM推荐的图像优先排序,即问题优先悖论。研究发现该悖论源于VLM计算两阶段的冲突。问题前置能引导感知,但后续答案生成阶段难以关注到问题。解决方法是问题回声,即在图像两侧重复问题,一侧引导感知,另一侧供答案生成阶段读取。实验表明,回声提示解决了悖论,在多个基准测试中超越最佳单通道排序且无需训练等。

英文摘要

Where should the question go in a vision-language model (VLM) prompt: before the image or after it? Intuition says before: knowing what is asked should tell the model where to look. Yet across visual question answering benchmarks, question-first prompting consistently underperforms the image-first ordering recommended for frontier VLMs, a phenomenon we term the question-first paradox. We trace this paradox to a conflict between two stages of VLM computation. Logit-lens and attention probes show that question-first prompting steers perception, shifting image patch representations toward question-relevant concepts. But downstream, stranded behind hundreds of image tokens, the question is barely attended by the answer token, which instead commits to image-driven, often wrong answers. Causal attention knockout confirms that the answer reads the question only when it follows the image. This diagnosis yields a training-free fix: question echoing, restating the question on both sides of the image so one copy steers perception while the other is available at answer time. A similar division of labor appears in a fifty-year-old finding on human 'adjunct questions', where repeating a question before and after a passage improves comprehension. Echoing the image as well brings further gains by restoring the whole-image view otherwise lost by a causal decoder. The paradox holds across five open VLMs, costing up to 17.5 group-accuracy points. Echoed prompts recover most of the gap and, on NaturalBench and Winoground, surpass the best single-pass ordering by up to 19 group-accuracy points on Winoground, with no training, fine-tuning, or architecture change. The paradox reveals a tension between steering what a model sees and preserving access to what it was asked; echoing resolves this through prompt design. Project Page: https://rakshanda-cmu.github.io/ask-twice-look-twice/

CommentsAccepted at the eXCV Workshop, ECCV 2026. Project page: https://rakshanda-cmu.github.io/ask-twice-look-twice/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑