发表机构
Pázmány Péter Catholic University; Chulalongkorn University(帕兹马尼·彼得天主教大学; 朱拉隆功大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对Sen1Floods11和GEOID-Flood数据集,通过多种受控实验评估地表水分割模型的排名稳定性、输入依赖等,发现聚合指标排名存在局限,需多维度证据支撑评估结论。
AI 中文摘要
地表水分割的性能评估通常采用全局交并比(IoU)等聚合指标对模型配置进行排名。然而,配置排名本身无法说明某一系统表现更优的原因、相近排序是否稳定,也无法体现预测对单个输入的依赖程度。本研究主要针对Sen1Floods11数据集,通过重复配置比较、配对测试芯片分析、固定检查点输入压力测试及地理加权,结合对GEOID-Flood数据集上监督输入配置的针对性二次评估,考察上述差异。跨模态学生模型在Sen1Floods11上取得最高的3个随机种子平均IoU,但相近排序随种子和地理加权而变化,且Swin-UNet与U-Net的辅助输入排名存在差异。GEOID-Flood评估显示,监督辅助输入的效应存在显著一致性,不过具体架构排序仍依赖配置。固定检查点测试进一步证实模型对地形和WorldCover的依赖,但未证实干净输入的性能优势,同时目标语义及后续WorldCover先验将评估限制于回顾性全水体分割。这些结果表明,聚合指标对完整配置排名仍有用,但排名稳定性、组件归因、输入依赖及部署范围需不同证据,性能评估应使报告的证据与所提出的主张匹配。
英文摘要
Performance evaluation for surface-water segmentation commonly uses an aggregate metric such as global intersection-over-union (IoU) to rank model configurations. However, a configuration ranking does not by itself establish why one system performs better, whether a close ordering is stable, or how strongly predictions rely on individual inputs. We examine these distinctions primarily on Sen1Floods11 through repeated configuration comparisons, paired test-chip analysis, fixed-checkpoint input stress tests, and geographic reweighting, with a targeted secondary evaluation of supervised input configurations on GEOID-Flood. The cross-modal student achieves the highest three-seed mean IoU on Sen1Floods11, but close orderings vary across seeds and geographic weighting, while ancillary-input rankings differ between Swin-UNet and U-Net. The GEOID-Flood evaluation shows substantial agreement in supervised ancillary-input effects, although the exact architecture ordering remains configuration dependent. Fixed-checkpoint tests further establish reliance on terrain and WorldCover without establishing a clean-input performance benefit, while target semantics and the later WorldCover prior restrict the evaluation to retrospective all-water segmentation. These results show that aggregate metrics remain useful for ranking complete configurations, but ranking stability, component attribution, input reliance, and deployment scope require distinct evidence. Performance evaluation should therefore match the evidence reported to the claim being made.