看起来相同,答案却不同:面向稳健视觉-语言推理的翻转方向引导
Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning
浏览论文内容
中文总结 AI 辅助
针对视觉-语言模型在细微图像变化下答案翻转的问题,提出无需训练的FlipDir方法,通过低秩子空间引导和边际门控稳定推理,并引入VisFlip基准验证,在18个设置中优于现有方法。
中文摘要 AI 辅助
视觉-语言模型(VLMs)在视觉推理任务上表现出色,然而,即使图像看起来几乎相同,常规图像采集和处理过程中的细微变化也可能改变其推理轨迹。在长程生成中,由此产生的激活偏移可能在解码步骤中累积,逐步改变推理令牌,并最终改变最终答案,这一现象被称为答案翻转。为解决这一不稳定性,我们提出了FlipDir(翻转方向引导),一种无需训练、推理时的方法,该方法从原始输入与导致答案翻转的输入之间的对比对中估计一个低秩的翻转诱导激活子空间,并在解码过程中选择性地引导隐藏状态。一个基于边际的门控将子空间衰减限制在不确定的解码步骤,从而恢复原始预测,同时保持稳定的预测不变。为了在固定测试集上的准确性或一致性之外评估稳健性,我们引入了VisFlip,一个基准框架,该框架为目标模型和视觉变化设置构建评估组,以分别评估原始预测的恢复和稳定预测的保持。VisFlip涵盖九个数据集-变化组合,涉及科学推理、机器人场景理解和医学视觉问答,覆盖了各领域常见的细微视觉变化。在18个设置上的实验表明,FlipDir在恢复与保持的联合指标上持续优于现有方法。我们将公开我们的代码。
英文摘要
Vision-language models (VLMs) achieve strong visual reasoning performance, yet subtle changes from routine image capture and processing can alter their reasoning trajectories even when images appear nearly identical. In long-horizon generation, the resulting activation shifts may accumulate across decoding steps, progressively altering reasoning tokens and ultimately changing the final answer, a phenomenon referred to as answer flips. To address this instability, we propose FlipDir (Flip-Direction Steering), a training-free inference-time method that estimates a low-rank flip-inducing activation subspace from contrastive pairs of original and answer-flipping inputs and selectively steers hidden states during decoding. A margin-based gate limits subspace attenuation to uncertain decoding steps, recovering original predictions while preserving stable ones. To evaluate robustness beyond accuracy or consistency on fixed test sets, we introduce VisFlip, a benchmark framework that constructs evaluation groups for a target model and visual variation setting to separately assess recovery of original predictions and preservation of stable ones. VisFlip spans nine dataset-variation combinations across scientific reasoning, robot-scene understanding, and medical VQA, covering subtle visual variations common in each domain. Experiments across 18 settings demonstrate that FlipDir consistently outperforms existing methods on the combined recovery and preservation metric. We will make our code publicly available.
发表机构
- KAIST(韩国科学技术院)
- NAVER Cloud
- The Ohio State University(俄亥俄州立大学)
- Adobe Research(Adobe研究院)
- AITRICS
机构由 AI 辅助整理,请以论文原文为准。