arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

只需持续提示:评估视觉语言模型中的重复苏格拉底式提示

Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs

Shayda Moezzi, Bishoy Galoaa, Lorena Genua, Taskin Padir, Sarah Ostadabbas

arXiv 2607.14099首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言模型在反复提示下的稳定性,引入JKP多轮评估框架,用三种策略对模型提问,在STAR基准子集上评估GPT-4o、Gemini 2.5 Pro和Qwen3-VL-30B,发现重复提示利弊兼具且因模型而异,揭示了模型压力反应概况。

AI 中文摘要

在现实世界中部署视觉语言模型(VLM)不仅需要强大的视觉推理能力,还需要在持续对话压力下保持稳定。我们引入了Just Keep Prompting(JKP),这是一个多轮评估框架,用于衡量当用户反复挑战、质疑或反驳模型答案时VLM的认知稳定性。JKP使用三种策略对模型进行多达10轮的后续提问:对抗性否定(重复拒绝)、纯粹的苏格拉底式询问(重复要求重新评估确定性)和上下文感知苏格拉底式总结(在要求重新考虑之前反映模型先前的理由)。我们在STAR基准的一个子集上对GPT-4o、Gemini 2.5 Pro和Qwen3-VL-30B进行了720次多轮运行的评估。从第0轮到第10轮,总体准确率变化不大,但轨迹级分析显示出显著的不稳定性:正确答案倒退,错误答案恢复,许多运行显示出答案反复翻转。重复提示的好处有限,而且往往起到破坏稳定的作用,而不是推理辅助作用。这种影响强烈依赖于模型:Qwen3-VL-30B最终准确率最高,但在直接矛盾下会自信地给出错误答案;Gemini 2.5 Pro相对稳定,但token成本高;GPT-4o最脆弱且波动大。这些发现表明,多轮VLM评估不仅捕捉了额外的推理,还捕捉了压力反应概况:模型在反复挑战下如何权衡视觉基础、校准和对话合规性。

英文摘要

Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversational pressure. We introduce Just Keep Prompting (JKP), a multi-turn evaluation framework that measures VLM epistemic stability when users repeatedly challenge, question, or contradict a model's answer. JKP probes models for up to 10 follow-up turns using three strategies: Adversarial Negation (repeated rejection), Pure Socratic Interrogation (repeated calls to reassess certainty), and Context-Aware Socratic Summarization (reflecting the model's prior rationale back before asking for reconsideration). We evaluate GPT-4o, Gemini 2.5 Pro, and Qwen3-VL-30B on a subset of the STAR benchmark across 720 multi-turn runs. Aggregate accuracy changes modestly from Turn 0 to Turn 10, but trajectory-level analysis reveals substantial instability: correct answers regress, wrong answers recover, and many runs exhibit repeated answer flipping. Repeated prompting has bounded upside and often acts as a destabilizer rather than a reasoning aid. The effect is strongly model-dependent: Qwen3-VL-30B achieves the highest final accuracy but becomes confidently wrong under direct contradiction; Gemini 2.5 Pro is comparatively stable but token-expensive; GPT-4o is the most brittle and oscillatory. These findings reveal that multi-turn VLM evaluation captures not just additional reasoning but pressure-response profiles: how models trade off visual grounding, calibration, and conversational compliance under repeated challenge.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑