arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25993cs.CV

超越缩放:学习用于超高分辨率遥感的多工具视觉推理

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

Fengxiang Wang, Jiangnan Huang, Mingshuo Chen, Yueying Li, Yang Shi, Junwei Luo, Haoyu Wang, Yansheng Li, Jing Zhang, Haiyan Zhao, Wenjing Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对超高分辨率遥感图像给多模态大语言模型带来的挑战,提出GeoMTVR数据集,结合监督微调与以工具注意力为重点的强化学习算法开发GeoLens,实验证明其在多工具视觉推理方面优于直接推理和单工具放大基线。

中文摘要 AI 辅助

超高分辨率(UHR)遥感图像在城市尺度场景中提供了细粒度的地球观测证据,但对多模态大语言模型(MLLM)构成了根本挑战,因为任务相关证据往往稀疏、局部且在极大的视觉背景中空间分散。自然的解决办法是为MLLM配备放大工具进行主动局部检查。然而,通过对XLRS-Bench的初步研究发现,放大仅部分有效。在此基础上,引入了GeoMTVR数据集,它包含13K UHR VQA样本。还提出了一种以工具注意力为重点的强化学习算法。结合在GeoMTVR上的监督微调与该强化学习算法,开发了GeoLens。实验表明,GeoLens始终优于直接推理和单工具放大基线。

英文摘要

Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental challenge for multimodal large language models (MLLMs): task-relevant evidence is often sparse, local, and spatially dispersed across extremely large visual contexts. A natural solution is to equip MLLMs with zoom-in tools for active local inspection. However, through a pilot study on XLRS-Bench, we find that zoom-in is only partially effective: it resolves easy and medium-level tasks with locally recoverable evidence, but saturates on hard cases requiring global search, multi-region comparison, path planning, or dispersed-evidence reasoning. Motivated by this finding, we move beyond single-tool zoom-in and introduce GeoMTVR, a large-scale Geospatial Multi-Tool Visual Reasoning dataset built from wide-area satellite imagery. GeoMTVR contains 13K UHR VQA samples with interleaved reasoning trajectories, diverse visual tool calls, and returned visual observations, enabling models to learn question decomposition, tool selection, regional inspection, object-level grounding, auxiliary visual reasoning, and cross-tool evidence integration. Beyond supervised fine-tuning, we propose a tool-attention-focused reinforcement learning algorithm that concentrates optimization on critical tool-use decisions, including when to invoke tools, which tool to select, where to apply it, and how to interpret tool outputs. By combining SFT on GeoMTVR with our RL algorithm, we develop GeoLens, a multi-tool visual reasoning MLLM for UHR RS. Experiments show that GeoLens consistently outperforms direct reasoning and single-tool zoom-in baselines, achieving stronger accuracy, better evidence grounding, and more efficient tool-use trajectories.

发表机构

  • National University of Defense Technology(国防科技大学)
  • Wuhan University(武汉大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑