发表机构
School of Frontier Sciences, Nanjing University; Computer Network Information Center, Chinese Academy of Sciences(南京大学前沿科学学院; 中国科学院计算机网络信息中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对UHR遥感图像理解,提出免训练框架WeaveEarth,通过全局感知证据构建选最小支持证据集,经结构化证据推理增强VLM全局-局部联合推理能力,在多基准测试中表现优于现有方法。
AI 中文摘要
超高分辨率(UHR)遥感图像理解要求视觉语言模型(VLM)在有限计算预算下捕捉全局场景布局和稀疏但关键的局部细节。现有方法主要有被动感知和主动感知两种范式,但都存在不足。本文提出WeaveEarth,一个免训练框架,将UHR理解重新表述为全局上下文约束下的结构化证据构建与推理问题。具体包括全局感知证据构建以选择最小支持证据集,以及结构化证据推理将局部证据等编织成统一推理接口,增强VLM全局-局部联合推理能力。实验表明WeaveEarth在多个基准测试中优于现有方法。
英文摘要
Ultra-High-Resolution (UHR) remote sensing image understanding requires Vision-Language Models (VLMs) to capture both the global scene layout and sparse yet task-critical local details under limited computational budgets. Existing methods mainly follow two paradigms. One is passive perception, which relies on resolution expansion or token compression and may therefore discard fine-grained details. The other is active perception, which depends on multi-round zooming and search, but suffers from high latency, contextual fragmentation, and error accumulation. We argue that a more effective path toward UHR understanding lies not in accessing more, but in organizing better. To this end, we propose WeaveEarth, a training-free framework that reformulates UHR understanding as a problem of structured evidence construction and reasoning under global context constraints. Specifically, WeaveEarth first employs Global-Aware Evidence Construction to select a compact, low-redundancy, and spatially complementary Minimal Support Evidence Set. It then introduces Structured Evidence Reasoning, which weaves local evidence, spatial metadata, and relative topology into a unified reasoning interface, thereby enhancing the VLM's ability to perform global-local joint reasoning. Extensive experiments show that WeaveEarth consistently outperforms strong baselines and existing UHR methods across multiple UHR remote sensing benchmarks and multiple frozen VLM backbones. Code is available at https://github.com/XianZhi-Ma/WeaveEarth.