arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30959cs.CVcs.AI

LOCI:带精化循环的定位器-评判器

LOCI: A Locator-Critic with Refinement Loop

Walid Bousselham, Mathilde Caron, Arsha Nagrani, Cordelia Schmid

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉语言模型无法定位图像关键细节的问题,提出无需训练的LOCI框架,通过定位器与评判器的迭代精化循环实现自修正,在多个复杂视觉基准上取得当前最优结果,显著提升了开源与专有VLMs的准确率。

中文摘要 AI 辅助

视觉语言模型(VLMs)在需要复杂视觉理解的任务上仍然表现不佳。我们认为核心问题并非高层推理能力,而是无法定位图像中的关键细节。由于这一缺陷,VLMs常常基于有缺陷的感知依据生成看似合理但实则错误的推理。为解决该问题,我们提出LOCator-Critic(LOCI),这是一个无需训练的框架,将视觉搜索与证据验证解耦。LOCI使用一个定位器智能体提出候选视觉证据,另一个独立的评判器智能体评估该证据的相关性与充分性。这些智能体参与迭代精化循环,逐步优化证据,直至其足以回答给定问题。这种解耦的自修正过程带来了显著的性能提升,在多个复杂视觉基准上取得了当前最优结果。LOCI提升了两类模型的准确率:开源权重模型如Qwen3-VL(在V*上提升12.1,在HR-Bench上提升5.8,在VisualProbe-Hard上提升11.2),以及专有模型如Gemini 2.5 Pro(在V*上提升8.9,在HR-Bench上提升4.3,在VisualProbe-Hard上提升4.8)。

英文摘要

Vision-Language Models (VLMs) still struggle on tasks requiring complex visual understanding. We argue that the core issue is not high-level reasoning, but instead failing to locate critical details in the image. Due to this shortcoming, VLMs generate often plausible but incorrect reasoning based on flawed perceptual grounding. To address this, we propose Locator-Critic (LOCI), a training-free framework that decouples visual search from evidence verification. LOCI employs a Locator agent to propose candidate visual evidence and a separate Critic agent to evaluate its relevance and sufficiency. These agents engage in an iterative refinement loop, progressively improving the evidence until it is adequate to answer the given question. This decoupled, self-correcting process yields substantial performance gains, achieving state-of-the-art results on multiple complex visual benchmarks. LOCI improves accuracy for both open-weight models like Qwen3-VL (+12.1 on V*, +5.8 on HR-Bench and +11.2 on VisualProbe-Hard) and proprietary models like Gemini 2.5 Pro (+8.9 on V*, +4.3 on HR-Bench, +4.8 on VisualProbe-Hard).

发表机构

  • Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

↑