Hi-OPD:面向遥感图像的层级感知开放提示检测
Hi-OPD: Hierarchy-Aware Open-Prompt Detection for Remote Sensing Images
浏览论文内容
中文总结 AI 辅助
Hi-OPD 提出层级感知开放提示检测方法,通过层级安全负采样和路径多正样本监督修复多源遥感标注中祖先查询下的漏检问题,在多个基准上显著提升父类与祖父类 AP50。
中文摘要 AI 辅助
Hi-OPD 针对扁平开放提示训练中一个未被控制的失效模式:当多源遥感标注存在粒度不一致和缺失标签时,后代检索在祖先查询下不必持续存在。检测器可以在原子提示下定位轿车和货车,但在车辆提示下却会漏检同一实例;扁平 AP 无法暴露这种跨层级的不一致性。我们提出 Hi-OPD,一种层级感知的开放提示检测器,并从 175,644 条保留的训练图像/瓦片记录和映射到 153 个原子类别(具有稀疏层级和别名关系)的 348 万边界框中构建了 RS153-HierOPD。Hi-OPD 通过层级安全负采样、路径多正样本监督和单向向上一致性来学习祖先检索,同时通过逐源风险排除来处理可能缺失的标签。ConvVPE 利用检测器原生特征和共享对比头将 K-shot 支持边界框转换为文本兼容的嵌入。在 Track A 上,Hi-OPD 在 DIOR/DOTA-v2.0 上获得 79.7/72.3 的 AP50,高于文献报道的 OpenRSD 结果 76.7/71.8。在原始转换标注上的受控训练下,完整层级方案将 DOTA-v2.0 父类 AP50 从 7.2 提升至 71.5,将 FAIR1M 祖父类 AP50 从 31.6 提升至 71.4,而 DOTA-v2.0 原子类 AP50 从 71.4 变为 72.3。文本路径在三个常见源上达到 99.7% 的 CAR50(0.3% 违规),在 FAIR1M 祖父类关系上达到 99.9%/0.1%。在保留的 VEDAI 上,文本 AP50 为 75.9,比 OpenRSD 高 6.2 个点。联合 AP 和 CAR 表明,显式层级训练修复了该失效模式,同时保留了原子检测和提示迁移能力。
英文摘要
Hi-OPD addresses a failure mode left uncontrolled by flat open-prompt training: descendant retrieval need not persist under ancestor queries when multi-source remote sensing annotations exhibit inconsistent granularity and missing labels. A detector may localize \textit{car} and \textit{van} under atomic prompts yet miss the same instances under \textit{vehicle}; flat AP does not expose this cross-level inconsistency. We propose Hi-OPD, a hierarchy-aware open-prompt detector, and construct RS153-HierOPD from 175,644 retained training image/tile records and 3.48M boxes mapped to 153 atomic categories with sparse hierarchy and alias relations. Hi-OPD learns ancestor retrieval through hierarchy-safe negative sampling, path multi-positive supervision, and one-way upward consistency, while per-source risk exclusion handles potentially missing labels. ConvVPE converts K-shot support boxes into text-compatible embeddings using detector-native features and the shared contrastive head. On Track A, Hi-OPD obtains 79.7/72.3 AP50 on DIOR/DOTA-v2.0, above the literature-reported OpenRSD results of 76.7/71.8. Under controlled training on the original converted annotations, the full hierarchy recipe raises DOTA-v2.0 parent AP50 from 7.2 to 71.5 and FAIR1M grandparent AP50 from 31.6 to 71.4, while DOTA-v2.0 atomic AP50 changes from 71.4 to 72.3. The text path reaches 99.7% CAR50 (0.3% violation) across the three common sources and 99.9%/0.1% on FAIR1M grandparent relations. On held-out VEDAI, text AP50 is 75.9, 6.2 points above OpenRSD. Joint AP and CAR show that explicit hierarchy training repairs this failure mode while retaining atomic detection and prompt transfer.
发表机构
- Wuhan University(武汉大学)
- Institute of Seismology, China Earthquake Administration(中国地震局地震研究所)
机构由 AI 辅助整理,请以论文原文为准。