发表机构
TelePIX(TelePIX)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究少样本开放词汇遥感分割问题,核心方法是在冻结模型上通过文本反转从少量示例中恢复地址,主要贡献是提升了受影响类别的平均交并比,优于其他少样本方法。
AI 中文摘要
开放词汇分割可根据文本查询对任意类别进行标注,无需逐类训练,但在遥感图像上,其在某些类别上的表现不如在其他场景可靠。研究发现,差距主要源于文本查询而非分割模型。由于这些模型并非专门针对航空图像,用作查询的类名在视觉语言嵌入空间中往往是一个较弱的地址。通过在冻结模型上进行文本反转,仅保留推理文本,从少量示例中恢复该地址。在代表性基准上,受影响类别的平均交并比从3.9提高到39.4,在八个遥感数据集上优于在推理时注入视觉提示的少样本方法。
英文摘要
Open-vocabulary segmentation labels arbitrary categories from a text query without per-class training, yet on remote sensing imagery it underperforms on categories it handles reliably elsewhere. We find that much of this gap traces to the text query rather than to the segmentation model. Because these models are not specialized for overhead imagery, the class name that serves as the query is often a weak address into the vision-language embedding space. We show that a better name repairs part of the gap, while the remaining failures call for an address that the tested natural-language rephrasings do not provide. We recover that address from a few examples through textual inversion on a frozen model, keeping inference text only. On a representative benchmark this raises the mean intersection over union on the affected categories from 3.9 to 39.4, and across eight remote sensing datasets it improves over few-shot methods that instead inject visual prompts at inference.