arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向空间数据的基于设计的预测驱动推理

Design-Based Prediction-Powered Inference for Spatial Data

Shinichiro SHirota

arXiv 2608.10356首次发表:更新:

AI 中文总结

本文针对空间数据将预测驱动推理重构为基于设计的框架,推导了相关方差、阈值与双重稳健性结果,在48175单元总体及爱沙尼亚LUCAS应用中验证了设计匹配调优可提升精度。

AI 中文摘要

预测驱动推理(PPI)将全覆盖预测图与少量金标准样本相结合,无论预测图质量如何都能提供有效的置信区间。经典PPI理论从独立同分布(i.i.d.)标注出发,而空间标注通过调查设计或协变量驱动机制获取,且图误差可能存在空间相关性。我们将PPI重构为基于设计的框架:估计目标是固定空间总体的普查参数,随机性源于标注机制。我们推导了简单抽样和分层抽样下的精确设计方差、分块空间平衡发挥作用的阈值,以及选择依赖于图时用于估计倾向的三明治推断。核心结果涉及双重稳健性:当倾向设定错误时,正确的结果模型可确保超总体识别,但在实现的总体条件下会留下阶为σᵤ/√Nₑբբ,ᵥ的余项,其中Nₑբբ,ᵥ是权重所关注残差斑块的有效计数。在比率稳定标注下,该余项与标注数量无关,因此随着标注积累,覆盖率可能恶化。对于独立同分布或可交换残差场,Nₑբբ,ᵥ阶为N,当n/N→0时,余项在抽样误差旁可忽略;而空间相干依赖则会使余项产生约束。我们在包含48175个单元的完全枚举总体上重现了这一结果。爱沙尼亚LUCAS应用显示,功率调优和依赖诊断必须遵循设计:独立同分布PPI++调优会降低最优图的精度,而与设计匹配的调优可将标准误差降低约10%,并在7个土地覆盖估计目标上优于或匹配PPI++。合并残差诊断同样可能将空间结构化的层间变异误认为残差依赖。

英文摘要

Prediction-powered inference (PPI) combines a wall-to-wall prediction map with a small gold-standard sample to give confidence intervals valid whatever the map's quality. Canonical PPI theory starts from i.i.d.\ labelling, whereas spatial labels arrive through survey designs or covariate-driven mechanisms, and map errors may be spatially correlated. We recast PPI in a design-based framework: the estimand is a census parameter of a fixed spatial population, with randomness arising from the labelling mechanism. We derive exact design variances under simple and stratified sampling, a threshold for when blocked spatial balance pays, and sandwich inference for estimated propensities when selection depends on the map. Our main result concerns double robustness. With a misspecified propensity, a correct outcome model secures superpopulation identification but, conditional on the realised population, leaves a remainder of order $σ_u/\sqrt{N_{\mathrm{eff},v}}$, an effective count of the residual patches the weights see. Under ratio-stable labelling this remainder is free of the label count, so coverage can deteriorate as labels accumulate. For i.i.d.\ or exchangeable residual fields $N_{\mathrm{eff},v}$ is of order $N$ and the remainder is negligible beside sampling error when $n/N \to 0$; spatially coherent dependence instead makes it bind. We reproduce this on a fully enumerated population of $48{,}175$ cells. Estonian LUCAS applications show that power tuning and dependence diagnostics must respect the design: i.i.d.\ PPI++ tuning worsens precision for the best map, whereas design-matched tuning cuts standard errors by about $10\%$ and matches or beats PPI++ across seven land-cover estimands. Pooled residual diagnostics can likewise mistake spatially structured between-stratum variation for residual dependence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑