arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29483cs.CVcs.CL

GeoAgent:通过具身导航评估视觉语言模型的地理定位能力

GeoAgent: Evaluating VLM Geolocalization Through Embodied Navigation

Arka Mukherjee, Soham Roy, Kartikeya Trivedi, Shreya Ghosh

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出基于街景导航的GeoAgent基准,发现VLMs在地理定位中难区分区域模式,智能体导航可提升准确率,但模型存在地区偏差且自我改进能力差,明确了具身地理定位的挑战

中文摘要 AI 辅助

现代视觉语言模型(Vision-Language Models, VLMs)在图像地理定位任务中的表现远超人类基线,该任务在灾害响应、开源情报(OSINT)验证和位置隐私保护中至关重要。然而,多数针对AI在该任务上行为的研究仍局限于基于静态图像的检索、分类与预测。我们认为,该任务的真实复现应涉及具身导航,即多模态智能体自主探索周围环境以收集观测信息,再提交预测结果。为此,我们提出GeoAgent,这是一个基于智能体环境的基准,要求智能体在街景环境中导航,通过序列推理优化地理定位结果。我们的分析显示,现代VLMs难以区分区域模式,但在国家和大洲级别的预测上表现良好;与基于静态图像的基线相比,智能体导航在既定指标下显著提升了准确率。我们还发现,前沿模型架构在发达/发展中地区背景下存在严重偏差,且在给定错误先验时自我改进能力较差。总体而言,本研究明确了具身导航与地理空间推理的挑战,并公开发布了代码和GeoAgent环境:this https URL

英文摘要

Modern Vision-Language Models (VLMs) perform well above the human baseline in image geolocalization, a task critically important in disaster response, OSINT verification, and location privacy. However, most efforts to study AI behavior on the task remain limited to static image-based retrieval, classification, and predictions. We argue that faithful recreation of the task should involve embodied navigation, where a multimodal agent autonomously explores its surroundings to gather observations before submitting a prediction. To this end, we introduce \textbf{GeoAgent}, an agentic environment-based benchmark that requires agents to navigate Street View environments to refine their geolocalization through sequential reasoning. Our analysis shows that modern VLMs struggle to discern regional patterns while succeeding at country- and continent-level predictions. When compared to static image-based baselines, agentic navigation significantly improves accuracy across established metrics. We also note severe bias in a developed/developing region context across frontier model architectures and poor self-improvement capabilities given incorrect priors. Overall, our work establishes the challenges of embodied navigation and geospatial reasoning. We publicly release our code and the GeoAgent environment: https://geoagent-benchmark.github.io

发表机构

  • KIIT Bhubaneswar(布巴内斯瓦尔KIIT大学)
  • IIT Bhubaneswar(布巴内斯瓦尔印度理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑