arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAVER:面向VLM变化推理中错误恢复的口头证据选择性审计

SAVER: Selective Auditing of Verbal Evidence for Error Recovery in VLM Change Reasoning

Youdi Li

arXiv 2608.22857首次发表:更新:

发表机构

Panasonic Connect Co., Ltd.(松下连接公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLMs视觉变化推理的表述失败问题,提出轻量规则方法SAVER,通过选择性审计VLM输出的口头证据触发结构化重提示,在CLEVR-Change等基准上显著提升准确率。

AI 中文摘要

视觉语言模型(VLMs)即便视觉编码器包含充足信息,也常无法完成视觉变化推理。我们观察到,正确的VLM输出往往包含支持所声称变化的明确口头证据(物体名称、颜色、空间位置),而错误输出常缺乏此类证据。我们提出SAVER(Selective Auditing of Verbal Evidence for Error Recovery,面向错误恢复的口头证据选择性审计),这是一种轻量的基于规则的方法,会解析VLM响应中的此类证据,仅在证据缺失或不一致时触发结构化重新提示。在三个变化检测基准和四个VLMs上,SAVER显著提升了因模型未能清晰表述所观察内容(表述失败)导致错误的任务准确率,在CLEVR-Change上的提升幅度达+25.8%。这些证据模式也可由大语言模型(LLM)单次调用生成,在CLEVR-Change上与人工调优的门控效果相当。 ablation实验证实,驱动性能提升的是证据门控,而非仅重新提示。

英文摘要

Vision-language models (VLMs) frequently fail at visual change reasoning, even when their vision encoders contain sufficient information. We observe that correct VLM outputs tend to contain explicit verbal evidence (object names, colors, spatial locations) that supports the claimed change, while incorrect outputs often lack such evidence. We propose SAVER (Selective Auditing of Verbal Evidence for Error Recovery), a lightweight, rule-based method that parses VLM responses for this evidence and triggers structured reprompting only when evidence is missing or inconsistent. Across three change detection benchmarks and four VLMs, SAVER significantly improves accuracy on tasks where errors stem from the model failing to articulate what it saw (expression failures), with gains up to +25.8% on CLEVR-Change. The evidence patterns can also be generated by an LLM in a single call, matching the hand-tuned gate on CLEVR-Change. Ablation experiments confirm that the evidence gate, not reprompting alone, drives the improvement.

Comments19 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑