arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

空间有偏采样下的半监督学习

Semi-Supervised Learning under Spatially Biased Sampling

Bright Wiredu Nuakoh, Francky Fouedjio, Stephen Bradshaw, Yaw Kwaafo Awuah-Mensah, Wei Hong Tan, Emet Arya, Ebenezer Afrifa-Yamoah

arXiv 2609.07982首次发表:更新:

发表机构

African Institute for Mathematical Sciences (AIMS); Edith Cowan University; Rio Tinto; Kaplan Business School Pty Ltd; Centre for Marine Ecosystems Research(非洲数学科学研究所(AIMS); 埃迪斯科文大学; 力拓集团; 卡普兰商业学校私人有限公司; 海洋生态系统研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对空间有偏采样下边际失配问题,提出核加权局部散度诊断法,发现SSL性能呈阈值式崩溃(约0.71-0.77),非平稳性独立致损,模型在无标签区过度自信。

AI 中文摘要

标准半监督学习(SSL)通常依赖于有标签和无标签数据共享共同的边际分布。当标签在空间有偏或偏好性地点选择下收集时,这一假设常被空间有偏采样机制所违反。我们将边际失配、空间自相关和空间非平稳性视为三种不同的机制,通过标签采样集中参数、空间长度尺度和非平稳性强度参数独立变化,并探究失配如何降低SSL性能,聚类和流形假设是否在失配下仍然成立,以及由此产生的失败如何被诊断。利用受控的合成框架以及PovertyMap-WILDS、加利福尼亚州房价、社会经济和美国空气质量监测数据集,我们在考虑空间自相关和非平稳性的同时,系统地变化失配程度。通过一系列分析(包括分段回归变点分析),我们表明在合成生成器中,SSL性能并非逐渐下降,而是在分布失配足够严重时,在约0.71至0.77之间表现出阈值式崩溃。我们进一步证明,空间非平稳性独立于边际失配导致性能损失,并且模型在标签可用区域之外变得日益过度自信。为支持实际部署,我们评估了几种分布散度度量作为可靠性指标,并引入了一种核加权局部散度度量,该度量比朴素的局部方法能更稳定地估计空间失配。这些发现为更好地记录将无标签空间数据纳入半监督学习工作流程的风险提供了经验证据和诊断工具。

英文摘要

Standard semi-supervised learning (SSL) typically relies on labelled and unlabelled data sharing a common marginal distribution. This assumption is often violated by biased spatial sampling mechanism, when labels are collected under spatially biased or preferential site selection. We treat this marginal mismatch, spatial autocorrelation, and spatial non-stationarity as three distinct mechanisms, varied independently via a labelled-sampling concentration parameter, a spatial length scale, and a non-stationarity strength parameter, and ask how mismatch degrades SSL, whether the cluster and manifold assumptions survive it, and how the resulting failure can be diagnosed. Using a controlled synthetic framework alongside PovertyMap-WILDS, California housing, socio-economic and US air quality monitoring datasets, we systematically vary the degree of mismatch while accounting for spatial autocorrelation and non-stationarity. Through a series of analyses including a segmented-regression changepoint, we show that in the synthetic generator, SSL performance does not degrade gradually but instead exhibits a threshold-like breakdown between approximately 0.71 and 0.77 once distribution mismatch becomes sufficiently severe. We further demonstrate that spatial non-stationarity contributes to performance loss independently of marginal mismatch and that models become increasingly overconfident outside the regions where labels are available. To support practical deployment, we evaluate several distribution-divergence measures as indicators of reliability and introduce a kernel-weighted local divergence metric that provides a more stable estimate of spatial mismatch than a naïve localised approach. These findings provide empirical evidence and diagnostic tools for better documenting the risk of incorporating unlabelled spatial data into semi-supervised learning workflows.

Comments26 pages, 9 figures, 9 tables. Supplementary material included as an ancillary file

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑