发表机构
Central South University; Jiangxi Normal University; Hohai University; Renmin University of China; The Chinese University of Hong Kong(中南大学; 江西师范大学; 河海大学; 中国人民大学; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GeoRefer-Bench提出可验证地理空间指代分割基准,用逻辑形式与精确查询成功率评估,揭示掩膜重叠不足以证明关系推理,最强模型EQS仅74.1。
AI 中文摘要
航拍影像中的指代分割本质上是关系性的:查询可能要求“道路以北的建筑物”或“最靠近住宅区的池塘”,因此正确的指代对象可以包含一个对象、多个对象或没有对象。现有基准主要对掩膜重叠进行评分,这无法验证模型是否真正解决了所陈述的空间关系。我们引入了GeoRefer-Bench,一个用于可验证地理空间指代分割的基准。每个查询由度量场景图上的可执行逻辑形式表示,预测通过精确查询成功率(EQS)进行评估,只有当返回的实例集合与查询所表示的集合完全匹配时,EQS才满足。GeoRefer-Bench包含700个完整的2048x2048无人机场景(2.94 Gpx),地面采样距离为12.5厘米和25厘米,共26,217个实例、142,796个空间关系和20,916个可执行查询,涵盖五个推理级别。它还包括每个查询的三个释义、24.0%的不可回答查询、2,477个反事实对和五个防泄漏评估分割。一项独立审计重新推导了对象几何、掩膜所有权、关系值、查询执行和分割来源,在全部700个场景中发现零问题。忽视关系的策略可以保留非平凡的mIoU,但总体EQS最多达到22.7,表明仅重叠并不能证明关系基础。在十五个当前模型中,最强的达到74.1 EQS,但从第1级的98.9下降到第5级的60.5,而十个模型在两跳查询上的EQS低于5。GeoRefer-Bench将地理空间指代分割从掩膜匹配转变为可验证的指代解析。
英文摘要
Referring segmentation in overhead imagery is inherently relational: a query may ask for the buildings north of the road or the pond closest to a residential area, so the correct referent can contain one object, several objects, or none. Existing benchmarks mainly score mask overlap, which cannot verify whether a model actually resolved the stated spatial relation. We introduce GeoRefer-Bench, a benchmark for verifiable geospatial referring segmentation. Each query is represented by an executable logical form over a metric scene graph, and predictions are evaluated with Exact Query Success (EQS), which is satisfied only when the returned instance set exactly matches the set denoted by the query. GeoRefer-Bench contains 700 whole 2048x2048 UAV scenes (2.94 Gpx) at 12.5 and 25 cm ground sampling distance, 26,217 instances, 142,796 spatial relations, and 20,916 executable queries spanning five reasoning levels. It further includes three paraphrases per query, 24.0% unanswerable queries, 2,477 counterfactual pairs, and five leakage-controlled evaluation splits. An independent audit re-derives object geometry, mask ownership, relation values, query execution, and split provenance, finding zero issues across all 700 scenes. Relation-blind strategies can retain non-trivial mIoU while achieving at most 22.7 EQS overall, showing that overlap alone does not certify relational grounding. Across fifteen current models, the strongest reaches 74.1 EQS but drops from 98.9 at level 1 to 60.5 at level 5, while ten models score below 5 EQS on two-hop queries. GeoRefer-Bench turns geospatial referring segmentation from mask matching into verifiable reference resolution.
Commentshave some mistakes