发表机构
University of Southern California; Amazon Inc.(南加州大学; 亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉语言模型难以理解否定语义的问题,提出免训练的骨架与策略提示方法,通过抽象问题结构、检索同骨架示例并生成回答策略,在多个否定VQA基准上取得最先进性能且无需参数更新。
AI 中文摘要
尽管视觉语言模型(VLMs)在广泛的视觉问答(VQA)任务上表现出强劲性能,但这些模型在面对包含否定从句的问题时,始终难以理解否定语义,并产生错误答案。为解决这一局限,我们提出骨架与策略提示(SSP),一种免训练的上下文学习方法,在不进行任何参数更新的情况下提升VLM的否定理解能力。给定一个否定问题,我们的方法首先将底层问题结构抽象为骨架,从轻量级问题池中检索一小批同骨架问题,然后提示VLM分析它们共有的否定模式,并综合成一句回答策略。该骨架与策略被前置到测试样本前,以引导模型正确应对否定问题。在多个否定VQA基准上的实验表明,SSP在聚焦否定的VQA任务上达到了最先进的性能,同时保持了计算效率。
英文摘要
Despite the strong performance of Vision-Language Models (VLMs) on a wide range of visual question answering (VQA) tasks, these models consistently struggle to understand negation and produce incorrect answers when questions involve negated clauses. To address this limitation, we propose Skeleton-and-Strategy Prompting (\textbf{SSP}), a training-free, in-context learning method that improves VLM negation understanding capabilities without any parameter updates. Given a negation question, our method first abstracts the underlying question structure into a skeleton, retrieves a small set of same-skeleton questions from a lightweight question pool, then prompts the VLM to analyze their shared negation pattern and synthesize a single-sentence answering strategy. The skeleton and strategy are prepended to the test sample to guide the model correctly tackle the negation problems. Experiments on multiple negation VQA benchmarks show that SSP achieves state-of-the-art performance on negation-focused VQA tasks while remaining computationally efficient.