arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31747cs.CVeess.IV

一瞥地球:面向超高分辨率遥感理解的免训练主动聚焦

The Earth in One Gaze: Training-Free Active Focus for UHR Remote Sensing Understanding

Yao Zhang, Pengyu Dai, Wei Guo, Jian Liang, Jian Song, Yafei Ou, Hongruixuan Chen, Naoto Yokoya

AI总结:

GazeEarth提出免训练主动聚焦框架,通过问题引导区域选择与整场景中央凹观察,在固定像素预算下提升UHR遥感理解,基准平均准确率提升4.6-9.4个百分点。

AI中文摘要:

多模态大语言模型(MLLMs)在有限的视觉输入预算内解释超高分辨率(UHR)遥感(RS)图像时,必须平衡局部细节与场景上下文。现有的基于选择的方法要么通过相关性评分修剪标记并选择补丁,要么通过重复检查进行主动裁剪。这两种策略都没有直接在连续的整场景视图中重新分配像素:前者保留选定的标记或补丁,后者重新编码与周围环境分离的裁剪区域。我们的初步研究发现,冻结的MLLM已经能产生有用的问题引导的空间请求,但对选定区域的基于裁剪的检查并不能持续改善其答案。因此,我们将UHR理解表述为在固定像素预算下何处花费的问题。基于此,我们引入了GazeEarth,一个简单而有效的免训练框架,将问题引导的区域选择与整场景中央凹观察相结合。MLLM从索引概览中选择证据单元;一个确定性的、保持拓扑的扭曲将原始图像重采样到固定大小的画布上,放大其共享邻域同时压缩外围;同一个冻结模型从这个聚焦视图中回答,最多使用两次MLLM调用,无需外部选择器或迭代搜索。在三个UHR遥感基准和四个冻结骨干网络上,GazeEarth相比直接回答将基准平均准确率提高了4.6到9.4个百分点,相比概览回答提高了3.4到4.3个百分点,优于任务训练的方法。我们的分析表明,现有MLLM可以自行指导在UHR图像中何处观看,并且它们从选定证据中推断出的内容取决于该证据的呈现方式。

英文摘要:

Multimodal large language models (MLLMs) must balance local detail against scene context when interpreting ultra-high-resolution (UHR) remote sensing (RS) imagery within a limited visual-input budget. Existing selection-based methods either prune tokens and select patches through relevance scoring, or crop actively through repeated inspection. Neither strategy directly redistributes pixels within a continuous full-scene view: the first retains selected tokens or patches, and the second re-encodes a crop detached from its surroundings. Our pilot study finds that a frozen MLLM already produces useful question-guided spatial requests, yet crop-based inspection of the selected regions does not consistently improve its answers. We therefore formulate UHR understanding as a question of where to spend a fixed pixel budget. Based on this, we introduce GazeEarth, a simple-yet-effective training-free framework that couples question-guided region selection with full-scene foveated observation. The MLLM selects evidence cells from an indexed overview; a deterministic, topology-preserving warp resamples the original image onto a fixed-size canvas, enlarging their shared neighborhood while compressing the periphery; the same frozen model answers from this focused view, using at most two MLLM calls and no external selector or iterative search. Across three UHR remote sensing benchmarks and four frozen backbones, GazeEarth improves benchmark-averaged accuracy by 4.6 to 9.4 percentage points over direct answering and 3.4 to 4.3 over overview answering, outperforming task-trained methods. Our analyses show that existing MLLMs can guide where to look in UHR images on their own, and that what they can infer from the selected evidence depends on how that evidence is presented.

↑