arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向跨空间时间序列的预测驱动推断

Prediction-powered inference for time series across space

Shahzar Rizvi, David Burt, Vishwak Srinivasan, Renato Berlinghieri, Stefano Del Col, Tamara Broderick

arXiv 2610.08715首次发表:更新:

发表机构

MIT LIDS; UMass Amherst; Stanford University; Assicurazioni Generali(麻省理工学院信息与决策系统实验室; 马萨诸塞大学阿默斯特分校; 斯坦福大学; 忠利保险)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对时空短标签序列与长未标签序列,提出预测驱动推断结合HAC的方法,提供可靠的点估计和置信区间,优于自然替代方法。

AI 中文摘要

以下主题在时空环境中很常见:我们观察到一段相对较短、较近时间内的协变量和标签对序列。我们可以在更长的时间段内获得未标记的协变量。数据在多个空间位置上被观测到。例如,作物产量可能在大地理区域内的最近几年被观测到,而天气数据(对作物产量有信息量)则可在更长的时间段内获得。目标是估计每个空间位置未来期望的标签(如作物产量),并为该值提供有效的置信区间。仅观测到的时间段太短,无法获得可靠的估计。用机器学习填补缺失标签可能导致显著偏差。预测驱动推断(PPI)可以纠正这种偏差,但它依赖于独立同分布(i.i.d.)假设,该假设在我们预期的时序依赖下不成立。异方差和自相关一致(HAC)程序考虑了时间相关性,但尚未适应部分标签被填补的情况。我们在以下条件下提供可靠的点估计和置信区间:短标记时间序列(跨空间位置)、较长的未标记时间序列,以及一个给定协变量预测标签的不完美预测器。我们展示了我们的方法优于自然的替代方法。

英文摘要

The following motif is common in spatiotemporal settings: we have a sequence of covariate and label pairs observed for a relatively short, recent time period. We have access to unlabeled covariates over a longer time period. Data is observed over many spatial locations. For instance, crop yield might be observed over a large geographical area for recent years, but weather data (which is informative about crop yield) is available for a much longer period. The goal is to estimate, at each spatial location, the expected label (e.g., crop yield) in the future and provide a valid confidence interval for this value. The observed time period alone is too short for reliable estimates. Imputing missing labels with machine learning can cause substantial bias. Prediction-powered inference (PPI) can correct for this bias, but it relies on an i.i.d. assumption that breaks under our expected temporal dependencies. Heteroskedasticity and autocorrelation consistent (HAC) procedures account for temporal correlation, but have not been adapted to cases where some labels are imputed. We provide reliable point estimates and confidence intervals given: short labeled time series (across spatial locations), a longer unlabeled time series, and an imperfect predictor of labels given covariates. We show our method outperforms natural alternatives.

CommentsAccepted to TS-LIMITS Workshop at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑