arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReGround:通过自诊断与视觉重检验恢复多步推理中的视觉接地

ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination

Lei Peng, Shuai Lv, Wei Hu

arXiv 2608.04385首次发表:更新:

发表机构

University of Science and Technology of China; School of Artificial Intelligence and Data Science; State Key Laboratory of Precision and Intelligent Chemistry(中国科学技术大学; 人工智能与数据科学学院; 精准与智能化学国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ReGround是无需架构修改或外部工具的两阶段框架,通过自诊断与视觉重检验解决VLMs多步推理中的视觉接地丢失问题,在八个基准上获一致增益且推理开销适度。

AI 中文摘要

视觉-语言模型(VLMs)在多步推理过程中常出现视觉接地丢失的问题:随着推理链变长,后续推理步骤越来越依赖语言先验而非图像证据。我们在四个基准的2510个重新检验样本中,识别出与这种性能下降相关的一致基准级特征:图像token的注意力熵通常在第1轮下降,在图像重注入后再次上升。但我们发现,有效的视觉重检验需要两个互补要素:图像重注入与针对性自诊断。若无针对性诊断,重检验甚至会损害性能;而准确的自诊断能带来显著增益——在关键基准上产生数个百分点的波动,表明诊断质量是我们的设置中重检验有益或有害的关键因素。我们提出ReGround,这是一个两阶段框架,无需架构修改或外部工具,即可教导VLMs自诊断接地故障并选择性重检验视觉证据。通过能力自举,来自同一模型族的更强变体仅在数据构建期间提供诊断支架,而策略模型在推理时学习自主诊断并保留大部分辅助增益。在两个VLM主干的八个基准上的实验显示出一致的增益,尤其是在视觉密集型多步推理任务上,且相对于工具增强基线仅产生适度的推理开销。项目页面:this https URL。代码:this https URL。

英文摘要

Vision-Language Models (VLMs) often lose visual grounding during multi-step reasoning: as reasoning chains grow longer, later inference steps rely increasingly on language priors rather than image evidence. We identify a consistent benchmark-level signature associated with this degradation: across 2,510 re-examined samples from four benchmarks, attention entropy over image tokens typically decreases during Round 1 and rises again after image re-injection. However, we find that effective visual re-examination requires two complementary ingredients: image re-injection and targeted self-diagnosis. Without targeted diagnosis, re-examination can even hurt performance, whereas accurate self-diagnosis yields substantial gains -- a swing of several points on key benchmarks, indicating that diagnostic quality is a key factor in whether re-examination helps or hurts in our setting. We present ReGround, a two-stage framework that teaches VLMs to self-diagnose grounding failures and selectively re-examine visual evidence, without architectural modifications or external tools. Through capability bootstrapping, a stronger variant from the same model family provides diagnostic scaffolding only during data construction, while the policy model learns to diagnose autonomously at inference time and retains most of the assisted gains. Experiments on eight benchmarks across two VLM backbones demonstrate consistent gains, especially on visually intensive multi-step reasoning tasks, while incurring only modest inference overhead relative to tool-augmented baselines. Project page: https://sespoir.github.io/reground-page/ . Code: https://github.com/sespoir/ReGround .

CommentsAccepted to ACM Multimedia 2026 (MM '26). 8 pages main text, 4 figures, plus appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑