arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01950stat.MEstat.AP

用于通用域上大型多元空间数据快速推理的双图Matérn-Whittle(BMW)过程

Bigraphical Matérn-Whittle (BMW) Processes for Fast Inference of Big Multivariate Spatial Data on General Domains

Debangan Dey, Alokesh Manna, Christopher J. Geoga

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出双图Matérn-Whittle(BMW)过程,解决通用域大型多元空间数据的联合建模问题,可高效扩展至千万级空间-变量对,在模拟及空间转录组学数据中均表现优异。

中文摘要 AI 辅助

大型空间数据集如今会在数千个位置记录多个相关变量,且这些位置往往位于欧氏距离无法准确表征邻近关系的域上。核心难点在于联合建模跨变量依赖关系的同时保留变量级别的解释性。我们提出双图Matérn-Whittle过程,这是一种多元高斯过程,通过两个图解决了上述问题:空间图通过图拉普拉斯算子的分数幂生成每个变量的Matérn结构,因此该过程在任意拓扑上均有效,且每个变量具有独立的范围、平滑度和振幅;有向无环变量图编码了科学结构,我们证明每条缺失边对应相应场之间的精确条件独立性。我们进一步证明算子行列式不涉及跨依赖系数,这使得无矩阵似然评估和变量图的贝叶斯学习可大规模处理。估计仅需稀疏矩阵-向量乘积,可扩展至数千万个空间-变量对。在模拟中,该方法准确恢复了参数和图,在模型误设下保持鲁棒性,并在非凸域上将保留的预测误差减半。在包含19809个细胞和1122个基因的空间转录组学研究中,在笔记本电脑上用75分钟完成拟合,通过学习到的基因图进行跨变量借用将保留的预测误差降低了50%至91%。此外,还探讨了诸如估计 nugget 方差的可实现效率等理论挑战。

英文摘要

Large spatial data sets now record many correlated variables at many thousands of locations, often on domains where Euclidean distance misrepresents proximity. The central difficulty is modelling the cross-variable dependence jointly while retaining variable-level interpretation. We introduce the bigraphical Mat'ern-Whittle process, a multivariate Gaussian process that resolves this with two graphs. A spatial graph generates the Mat'ern structure of each variable through a fractional power of a graph Laplacian, so the process is valid on any topology, with per-variable range, smoothness and amplitude. A directed acyclic variable graph encodes the scientific structure: we prove that each absent edge yields an exact conditional independence between the corresponding fields. We further prove that the operator determinant does not involve the cross-dependence coefficients, which keeps matrix-free likelihood evaluation and Bayesian learning of the variable graph tractable at scale. Estimation requires only sparse matrix-vector products and scales to tens of millions of space-variable pairs. In simulations the method recovered parameters and graphs accurately, remained robust under misspecification, and halved held-out prediction error on a non-convex domain. In a spatial transcriptomics section with 19,809 cells and 1,122 genes, fitted in 75 minutes on a laptop, borrowing across the learned gene graph reduced held-out prediction error by 50 to 91 percent. Theoretical challenges, such as the achievable efficiency of estimating the variance of the nugget, are also explored.

发表机构

  • Texas A&M University(德克萨斯农工大学)
  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑