arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29374cs.CV

思考、观察与修正:多模态大语言模型中基于不一致感知的视觉自我修正

Think, Look, and Revise: Inconsistency-Aware Visual Self-Correction in MLLMs

Yu Cheng, Arushi Goel, Hakan Bilen

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对工具增强型多模态推理中工具输出缺乏验证的问题,提出ReVISE框架,通过带监督反思的训练数据集与强化学习目标奖励,实现错误检测与修正,在多个基准上取得性能提升。

中文摘要 AI 辅助

工具增强型多模态推理将外部工具(如目标检测、深度估计)集成至多模态大语言模型(MLLMs)中,以解决复杂视觉任务中的感知瓶颈问题。然而,现有方法极少对工具输出进行验证,限制了模型检测工具故障并从中恢复的能力。我们提出ReVISE框架,为工具增强型推理的MLLMs配备验证与动态错误恢复能力。ReVISE包含两部分:一是精心构建的训练数据集,用于监督反思行为,使模型能够验证工具导出的证据、在出现视觉不匹配时重新表述查询,并在外部工具不可靠时回退至固有接地能力;二是基于强化学习的目标奖励,鼓励内部反思并惩罚空间错位。在多个基准上的实验表明,该方法相比现有方法取得了一致的性能提升,凸显了工具增强型多模态推理中错误检测与修正的重要性。

英文摘要

Tool-augmented multimodal reasoning integrates external tools (e.g., object detection, depth estimation) into multimodal large language models (MLLMs) to address perceptual bottlenecks in complex visual tasks. However, existing approaches rarely verify tool outputs, limiting their ability to detect and recover from tool failures. We propose ReVISE, a framework that equips MLLMs with verification and dynamic error recovery for tool-augmented reasoning. ReVISE introduces (1) a curated training dataset that supervises reflective behaviors, enabling models to validate tool-derived evidence, reformulate queries when visual mismatches arise, and fall back to intrinsic grounding when external tools are unreliable; and (2) a reinforcement learning based targeted rewards that encourage internal reflection and penalize spatial misalignment. Experiments on several benchmarks demonstrate consistent improvements over existing methods, highlighting the importance of error detection and correction in tool-augmented multimodal reasoning.

发表机构

  • University of Edinburgh(爱丁堡大学)
  • NVIDIA Research(英伟达研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑