arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27456cs.CV

UrbanGround:从局部感知到真实城市中的空间能动性

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

  • Shanghai Jiao Tong University(上海交通大学)
  • National University of Singapore(新加坡国立大学)
  • Meituan(美团)
  • The Chinese University of Hong Kong(香港中文大学)
  • Shanghai University(上海大学)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, … 展开作者

Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang

中文总结 AI 辅助

本文提出首个基于香港3D地理空间数据构建的UrbanGround沙盒,探究MLLM智能体在真实城市中从局部感知转化为可靠行动的程度,发现其长时间探索时错误累积的缺陷,为相关研究提供支撑。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)可解析街景,但城市能动性取决于智能体移动后局部证据是否仍有用。本文探究当前MLLM智能体在复杂真实尺度城市中,能将局部城市感知转化为可靠行动的程度。我们提出UrbanGround,首个基于全港3D地理空间数据构建的香港物理受限复现环境,可验证该问题。UrbanGround支持第一人称视角的闭环交互,并提供用于导航的交互式地图,智能体可直接进入3D城市并以第一人称视角探索。我们的分析围绕三个研究问题展开:首先测试智能体经主动观察后,能否充分定位局部场景以回答空间问题;接着探究当目的地更远且更不明确时,该定位是否支持导航;最后检验所得行为在路线可用性和行人移动变化下是否仍有效。当前MLLM智能体通常在视觉识别和短距离空间推理方面具备有用的基础能力,但定向和感知行人的移动仍不可靠。其核心缺陷出现在长时间探索中,局部能力无法组合为持续的目标导向行为,且错误会累积而无法得到有效修正。我们希望UrbanGround能支持更广泛的研究,探究当前MLLM智能体在复杂、开放式城市环境中可靠探索的程度。

英文摘要

Multimodal large language models (MLLMs) can interpret a street view, but reliable urban action depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a real-scale city. We propose UrbanGround, an urban sandbox built from Hong Kong's territory-wide 3D geospatial data. It combines the city's geographic structure with continuous, collision-constrained control through a shared evaluation interface. Agents use first-person observations and an interactive map to select actions across tasks ranging from local question answering to long-horizon navigation. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can gather and interpret local visual evidence to answer spatial questions. Then we ask whether these abilities support navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far MLLM agents can explore reliably in open-ended urban environments.

补充信息

相关深度报道

↑