发表机构
Jožef Stefan Institute; Jožef Stefan International Postgraduate School; Trinity University; Singidunum University(约热夫·斯泰凡研究所; 约热夫·斯泰凡国际研究生学院; 三一大学; 辛吉杜努姆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究采用fANOVA量化遥感多标签分类中各设计选择的贡献,分析7个数据集后发现不同规模数据集的关键设计因素存在差异,为模型选择提供了依据。
AI 中文摘要
针对遥感图像(RSI)多标签分类(MLC)的深度学习(DL)模型基准测试,通常会产生无法在评估数据集之外推广的排名。本研究不再局限于排名,而是采用功能方差分析(fANOVA)系统量化各设计选择及其交互对性能变异性的贡献。我们开展两项实证分析,分别覆盖48个和20个DL模型,涉及网络架构、微调策略、学习策略、初始化等设计选择。通过在7个MLC RSI数据集上应用fANOVA,我们构建了捕获设计选择敏感性轮廓的数据集元表示。对这些元表示的层次聚类显示,数据集会根据对设计决策的响应自然分组,其模式与数据集固有属性(如规模、空间分辨率、标签空间复杂度)密切相关。研究发现,对于大规模数据集,微调策略和架构是主导因素;在数据受限场景中,初始化起决定性作用;而在中等规模场景中,架构与学习策略的交互决定性能。
英文摘要
Benchmarking deep learning (DL) models for multi-label classification (MLC) of remote sensing images (RSI) typically yields rankings that do not generalize beyond the evaluated datasets. In this work, we move beyond rankings by employing functional analysis of variance (fANOVA) to systematically quantify the contributions of individual design choices and their interactions to performance variability. We conduct two empirical analyses covering 48 and 20 DL models, respectively, spanning design choices such as network architecture, fine-tuning strategy, learning strategy, and initialization. By applying fANOVA across seven MLC RSI datasets, we construct dataset meta-representations that capture design-choice sensitivity profiles. Hierarchical clustering of these meta-representations reveals that datasets naturally group according to how they respond to design decisions, with patterns strongly linked to intrinsic dataset properties such as scale, spatial resolution, and label space complexity. Our findings show that for large-scale datasets, fine-tuning strategy and architecture are dominant factors, while in data-limited regimes, initialization becomes decisive. For intermediate regimes, the interaction between architecture and learning strategy governs performance.
CommentsTo appear at Discovery Science 2026