AI 中文总结
针对遥感语义分割的标注缺陷问题,提出无参考指标 CMF,通过反事实视图与视觉-语言判断器审计真实掩码,在多数据集上验证其性能优于基线,可作为标注审计工具。
AI 中文摘要
语义分割模型的训练与评估均以人工绘制的掩码为依据,但遥感标注通常较为粗糙、不完整或存在对齐偏差;高重叠分数可能反映的是与不完善标签的一致性,而非对图像的保真度,由此产生评估悖论。我们提出对比掩码保真度(Contrastive Mask Fidelity, CMF),这是一种无需训练、无参考的指标,可直接根据图像证据对候选类别掩码进行评分。CMF 会为每个掩码构建保留和擦除的反事实视图,并让冻结的视觉-语言判断器判断类别证据是否集中在掩码内部且在外部不存在。我们在受控掩码损坏上验证了 CMF,随后使用基于 SegEarth-OV3 构建的无训练开放词汇探针 Seg-Probe 的候选掩码,对十个遥感基准中的 10731 个图像-类别对进行审计,Seg-Probe 在十个数据集中的九个上优于现有基线。审计揭示了系统性的、类别依赖的标注偏差:建筑物、道路和汽车等人造类别在 62%-85% 的对中更倾向于候选掩码,而模糊的土地覆盖类别则更常倾向于人工标注。在三名标注者共识的盲测中,CMF 与专家判断的匹配度达 81%,优于仅保留评分、模型置信度和训练后的标签质量基线。最后,保守的类别仲裁产生的监督信号比原始标注和匹配替换对照更能提升跨域迁移性能,使 CMF 成为可扩展的真实掩码审计工具,而非假设其绝对可靠。
英文摘要
Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox. We introduce Contrastive Mask Fidelity (CMF), a training-free, reference-free metric that scores competing class masks directly against image evidence. CMF composites keep and erase counterfactual views of each mask and asks a frozen vision-language judge whether class evidence is concentrated inside the mask and absent outside. We validate CMF on controlled mask corruptions, then audit 10,731 image-class pairs across ten remote-sensing benchmarks using candidate masks from Seg-Probe, a training-free open-vocabulary probe built on SegEarth-OV3 that outperforms prior baselines on nine of ten datasets. The audit reveals systematic, class-dependent annotation distortion: man-made classes such as buildings, roads, and cars favor the candidate mask on 62-85% of pairs, whereas ambiguous land cover more often favors human annotations. On a blinded three-annotator consensus, CMF matches expert judgment on 81% of pairs, exceeding keep-only scoring, model confidence, and a trained label-quality baseline. Finally, conservative class-wise arbitration yields supervision that improves cross-domain transfer over raw annotations and matched replacement controls, positioning CMF as a scalable tool for auditing ground truth rather than presuming it infallible.