arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GeoSelect:用于无训练的指称遥感图像分割的空间程序执行

GeoSelect: Spatial-Program Execution for Training-Free Referring Remote Sensing Image Segmentation

Yuhang Jiang, Guohui Deng, Miaozhong Xu, Chao Ruan, Jinling Zhao, Linsheng Huang

arXiv 2607.03869首次发表:更新:

发表机构

School of Internet, Anhui University; National Engineering Research Center for Agro-Ecological Big Data Analysis & Application, Anhui University; State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(安徽大学互联网学院; 安徽大学农业生态大数据分析与应用国家工程研究中心; 武汉大学测绘遥感信息工程国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究无训练的指称遥感图像分割,提出GeoSelect方法,将指称重构成类型化空间程序执行,由语言模型合成程序,经检查器和执行器运行,贡献是提升分割指标且可检查中间过程。

AI 中文摘要

指称遥感图像分割在航空图像中分离自然语言表达所命名的对象。现有无训练方法对空间、比较和顺序关系控制弱。我们提出GeoSelect,将指称重构为类型化空间程序执行。冻结的纯文本语言模型合成表达式,格式检查器接受程序,确定性执行器运行它。执行明确,中间过程可检查,可靠性阶梯降级失败程序。GeoSelect在测试中取得高指标。

英文摘要

Referring remote sensing image segmentation segments the object named by a natural-language expression in an aerial image. Existing training-free methods resolve the expression through implicit vision-language activations or region-text similarity, which gives weak control over the spatial, superlative, and ordinal relations that dominate aerial referring, such as the rightmost ship or the second court from the left. We propose GeoSelect, a training-free pipeline that reframes referring as the execution of a typed spatial program. A frozen, text-only language model synthesises the expression into a small domain-specific language, a well-formedness checker accepts the program, and a deterministic executor runs it. The central abstraction is a single scored candidate set type under which every operator composes: continuous geometric fields realise position and proximity, while discrete set and order operators add the extremum, ordinal, top-k, and relational constructions that fields alone cannot express. Execution is explicit, so every intermediate is inspectable, and a reliability ladder degrades any failing program to the field-only special case. GeoSelect achieves 58.86 mIoU on RRSIS-D test and 55.27 mIoU on RISBench test, more than twice the best prior training-free method on RRSIS-D, with no referring supervision and on a single GPU. Under a fixed detector and segmenter, explicit execution improves over implicit selectors under the same backbone; the best-box-IoU and outcome-partition diagnostics motivate complementary tests of proposal recall and program-path behaviour, with the program path the clearer priority on RISBench, and an exposure audit shows comparable accuracy on the audited unseen subset. Code and configurations are available at https://github.com/Avalon-S/GeoSelect.

CommentsAccepted version. Published in IEEE Transactions on Geoscience and Remote Sensing, DOI: 10.1109/TGRS.2026.3734378. 22 pages

Journal refIEEE Transactions on Geoscience and Remote Sensing, vol. 64, 2026, Art. no. 5642022

DOI:10.1109/TGRS.2026.3734378

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑