发表机构
University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明卫星图像的有效样本量为$\Theta(n^2/r^2)$,提出紧的泛化界,并论证空间交叉验证的合理性,在合成数据及三个传感器图像上验证。
AI 中文摘要
遥感影像的机器学习分类器通常被评估时假设每个像素都是独立样本。空间自相关违反了这一假设,因为相邻像素携带冗余信息,这会夸大样本量。一幅卫星图像实际上包含多少个独立样本?对于一幅$n \times n$的图像,其空间相关性在$r$个像素的范围内持续存在,有效样本量为$\Theta(n^2/r^2)$,而非$n^2$。我们证明了这是空间相关数据上分类器的有限样本上界,并通过匹配的下界表明该速率是紧的,且没有算法能做得更好。我们将结果扩展到具有方向相关性和空间变化相关结构的图像。我们的结果证明了空间交叉验证的合理性,因为与相关范围成比例的分离的块留出法能达到最优泛化保证,而随机留出法可能低估置信区间宽度,低估因子与$r$成正比。我们在合成数据和来自三个传感器(Landsat 8、Sentinel-2和Sentinel-1)的卫星图像瓦片上验证了该理论。
英文摘要
Machine learning classifiers for remote sensing imagery are typically evaluated as though every pixel were an independent sample. Spatial autocorrelation violates this assumption, since neighboring pixels carry redundant information which inflates sample sizes. How many independent samples does a satellite image actually contain? For an $n \times n$ image whose spatial correlation persists over a range of $r$ pixels, the effective sample size is $Θ(n^2/r^2)$, not $n^2$. We prove this as a finite-sample upper bound for classifiers on spatially correlated data, and show via a matching lower bound that the rate is tight, and no algorithm can do better. We extend the results to images with directional correlation and spatially varying correlation structure. Our result justifies spatial cross-validation since block holdout with separation proportional to the correlation range achieves optimal generalization guarantees, while random holdout can underestimate confidence interval widths by a factor proportional to $r$. We validate the theory on synthetic data and satellite image tiles from three sensors (Landsat 8, Sentinel-2, and Sentinel-1).
Journal refIEEE Transactions on Geoscience and Remote Sensing (2026) vol. 64