AI 中文总结
该研究针对空间数据的主成分分析方法的三大缺陷,提出空间正交因子模型与线性复杂度的MM-EM算法,成功应用于空间转录组学和遥感数据的空间分布推断。
AI 中文摘要
主成分分析常被应用于空间数据,以推断空间变异的潜在模式。这类分析广泛应用于空间转录组学和环境科学等领域,其中空间变异模式由基因表达因子或遥感时间序列测量值来表征。已有诸多方法被提出用于将空间信息融入概率主成分分析框架,但现有方法存在三个主要缺陷:其一,载荷矩阵非正交,对这些载荷进行后续正交化会破坏原始先验空间信息;其二,现有方法假设空间先验具有平稳性;其三,现有方法通常无法实现与空间位置数量相关的线性时间计算复杂度。为解决这些问题,我们首先直接用正交载荷对模型进行参数化,对于先验分布,我们推导了具有k个唯一奇异值和m-k个重复奇异值的奇异值分解变换的采样分布。随后,我们证明在该模型下,正交载荷的最大后验估计量为S + (1/n)Σ的特征分解,其中S为经验协方差矩阵,Σ为先验空间协方差。我们开发了一种EM内的极小化-最大化(minorization-maximization-within-EM,MM-EM)算法,其计算复杂度与空间位置数量呈线性关系。我们进一步扩展MM-EM算法以处理预留位置,并开发了一种用于优化非平稳先验协方差的验证策略。我们将所提方法用于推断人类大脑空间转录组学案例研究中方向特定长度尺度的空间分布,以及撒哈拉以南非洲的大陆尺度物候学案例研究。
英文摘要
Principal component analyses are often applied to spatial data towards inference on latent modes of spatial variation. These analyses are widespread across domains including spatial transcriptomics and environmental sciences, where the modes of spatial variation are represented by corresponding factors of gene expression or remotely sensed time series measurements. Many methods have been proposed for incorporating spatial information into a probabilistic PCA framework; however, there are three main drawbacks to currently available approaches. First, the loadings matrices are not orthogonal, and subsequent orthogonalization of those loadings corrupts the original prior spatial information. Furthermore, currently proposed methods assume stationarity in their spatial prior. Finally, current methods typically do not achieve linear-time computational complexity with respect to the number of spatial locations. To resolve these problems, we first parameterize the model directly with orthogonal loadings. For the prior distribution, we derive the sampling distribution of an SVD transformation with $k$ unique and $m-k$ repeated singular values. We then show under this model that the maximum a posteriori estimator for the orthogonal loadings is the eigendecomposition of $S + \frac{1}{n}Σ$, where $S$ is the empirical covariance matrix and $Σ$ is the prior spatial covariance. We develop a minorization-maximization-within-EM algorithm that is linear in computational complexity with respect to the number of spatial locations. We further extend our MM-EM algorithm to handle held-out locations and develop a validation strategy for optimizing the nonstationary prior covariance. Our methodology is used to infer the spatial distribution of direction-specific length scales in a human brain spatial transcriptomics case study, as well as a continental-scale phenology case study in sub-Saharan Africa.