arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2602.09463cs.AI

SpotAgent:通过代理推理在大视觉语言模型中实现视觉地理定位

SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic Reasoning

  • Peking University(北京大学)
  • StepFun

机构由 AI 辅助整理,请以论文原文为准。

Furong Jia, Ling Dai, Wenjin Deng, Fan Zhang, Chen Hu, Daxin Jiang, Yu Liu

更新

AI总结:

SpotAgent通过代理推理提升大视觉语言模型在地理定位中的表现,结合外部工具验证和强化学习优化,实现更精确和可靠的定位结果。

AI中文摘要:

大视觉语言模型(LVLMs)在地理定位任务中展现出了强大的推理能力,但在现实场景中,当视觉线索稀少、长尾分布且高度模糊时,它们往往表现不佳。以往的方法受限于内部知识,常常无法提供可验证的结果,在面对复杂证据时只能给出自信但缺乏依据的预测。为了解决这些挑战,我们提出了SpotAgent框架,该框架将地理定位正式化为一种代理推理过程,利用专家级推理来协同视觉解释与工具辅助验证。SpotAgent通过ReAct图式利用外部工具(例如网络搜索、地图)主动探索和验证视觉线索。我们引入了一个三阶段的后训练流水线,从监督微调(SFT)阶段开始,用于基本对齐,随后是利用多代理框架生成的高质量轨迹进行的代理冷启动阶段,旨在培养工具调用专长。接着,通过强化学习进一步优化模型的推理能力。我们提出了一种空间感知动态过滤策略,通过根据空间难度优先学习可学习样本来提高强化学习阶段的效率。在标准基准上的广泛实验表明,SpotAgent实现了最先进的性能,有效缓解了幻觉问题,同时提供了精确且可验证的地理定位。

英文摘要:

Large Vision-Language Models (LVLMs) have demonstrated strong reasoning capabilities in geo-localization, yet they often struggle in real-world scenarios where visual cues are sparse, long-tailed, and highly ambiguous. Previous approaches, bound by internal knowledge, often fail to provide verifiable results, yielding confident but ungrounded predictions when faced with confounded evidence. To address these challenges, we propose SpotAgent, a framework that formalizes geo-localization into an agentic reasoning process that leverages expert-level reasoning to synergize visual interpretation with tool-assisted verification. SpotAgent actively explores and verifies visual cues by leveraging external tools (e.g., web search, maps) through a ReAct diagram. We introduce a 3-stage post-training pipeline starting with a Supervised Fine-Tuning (SFT) stage for basic alignment, followed by an Agentic Cold Start phase utilizing high-quality trajectories synthesized via a Multi-Agent framework, aiming to instill tool-calling expertise. Subsequently, the model's reasoning capabilities are refined through Reinforcement Learning. We propose a Spatially-Aware Dynamic Filtering strategy to enhance the efficiency of the RL stage by prioritizing learnable samples based on spatial difficulty. Extensive experiments on standard benchmarks demonstrate that SpotAgent achieves state-of-the-art performance, effectively mitigating hallucinations while delivering precise and verifiable geo-localization.

↑