arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

经验贝叶斯谱部分池化跨相关任务

Empirical-Bayes spectral partial pooling across related tasks

Lorenzo Mauri

arXiv 2610.07284首次发表:更新:

发表机构

Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对相关但异质任务中谱估计不稳定或掩盖任务结构的问题,提出层次谱收缩(HSS)经验贝叶斯框架,部分池化任务特定谱估计器,提升估计精度和样本外性能。

AI 中文摘要

谱方法是高维统计和机器学习中的核心方法,是协方差估计、矩阵去噪、表示学习、聚类和潜变量建模等程序的基础。在本工作中,我们主要关注高维因子模型的谱估计器,其中数据矩阵的前导奇异向量用于估计潜在结构和协方差参数。然而,在许多现代应用中,数据是在相关但异质的任务、研究、领域或人群中收集的。当样本量相对于维度有限时,将此类估计器分别应用于每个任务可能导致不稳定的估计,而完全合并数据可能掩盖有意义的任务特定结构。我们引入了层次谱收缩(\ exttt{HSS}),这是一个可扩展的经验贝叶斯框架,用于部分池化任务特定的谱估计器。该方法将每个任务的前导谱方向正则化到数据自适应学习的公共基上,同时允许收缩量在不同任务和谱分量之间变化。所提出的估计器作为经验奇异向量的替代贝叶斯回归公式中的后验均值出现。对于因子模型,\ exttt{HSS} 产生组特定载荷空间和协方差矩阵的部分池化估计器,并可通过将其经验奇异向量替换为层次正则化对应向量,与不同的谱估计器结合使用。更广泛地,相同的构造为将谱估计器扩展到相关但异质的数据集集合提供了一种机制。我们在合成实验和多研究基因表达数据中展示了相对于单独、完全池化和共享子空间方法,在估计精度和样本外性能方面的显著改进。

英文摘要

Spectral methods are central to high-dimensional statistics and machine learning, underlying procedures for covariance estimation, matrix denoising, representation learning, clustering, and latent variable modeling. In this work, we focus primarily on spectral estimators for high-dimensional factor models, where leading singular vectors of the data matrix are used to estimate latent structure and covariance parameters. In many modern applications, however, data are collected across related but heterogeneous tasks, studies, domains, or populations. Applying such estimators separately to each task can lead to unstable estimates when sample sizes are limited relative to dimension, while completely pooling the data can obscure meaningful task-specific structure. We introduce Hierarchical Spectral Shrinkage (\texttt{HSS}), a scalable empirical-Bayes framework for partially pooling task-specific spectral estimators. The method regularizes the leading spectral directions of each task toward a data-adaptivelylearned common basis, while allowing the amount of shrinkage to vary across tasks and spectral components. The proposed estimator arises as the posterior mean in a surrogate Bayesian regression formulation of the empirical singular vectors. For factor models, \texttt{HSS} yields partially pooled estimators of the group-specific loading spaces and covariance matrices and can be combined with different spectral estimators by replacing their empirical singular vectors with hierarchically regularized counterparts. More broadly, the same construction provides a mechanism for extending spectral estimators to collections of related but heterogeneous datasets. We demonstrate substantial improvements in estimation accuracy and out-of-sample performance over separate, fully pooled, and shared-subspace approaches in synthetic experiments and multi-study gene-expression data.

Comments27 pages (including references), 32 pages with supplemental, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑