arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24378stat.MEq-bio.QMstat.APstat.COstat.ML

使用高斯和扩散伽马先验对高维时空数据进行贝叶斯特征提取

Bayesian Feature Extraction using Gaussian and Diffused-gamma Priors for High Dimensional Spatio-Temporal Data

Garrett Frady, Dipak K. Dey, Shariq Mohammed

首次发表
浏览论文内容

中文总结 AI 辅助

针对含稀疏结构及时空依赖性的高维数据,开发贝叶斯特征提取框架,采用高斯和扩散伽马先验诱导结构化稀疏,经MCMC计算后验,引入两阶段提取过程,以EEG案例研究验证,提升了稀疏特征恢复及可解释性,广泛适用于相关高维结构化问题。

中文摘要 AI 辅助

在许多科学领域中都会出现具有稀疏结构和时空依赖性的高维数据。我们为时空设置开发了一个贝叶斯特征提取框架,该框架采用高斯和扩散伽马先验来诱导结构化稀疏性。建模框架通过布雷格曼散度指定一般似然,从而与一系列损失函数和测量模型兼容。通过马尔可夫链蒙特卡罗(MCMC)进行后验计算,并基于后验样本引入两阶段特征提取过程,以稳定跨空间和时间的选择。我们通过多主体脑电图(EEG)案例研究来说明该方法,该研究考察慢性酒精暴露与不同脑区活动之间的关系。我们首先在每个时间点拟合二元分类模型,然后在两阶段特征提取管道中使用错误发现率控制的筛选和后续聚类来识别活跃脑区。分析表明,我们提出的先验改进了稀疏特征的恢复,并在存在时空依赖性的情况下增强了可解释性。该框架广泛适用于需要精确特征选择和推理的高维结构化问题。实现该模型的代码可通过GitHub公开获取。

英文摘要

High-dimensional data with sparse structure and spatio-temporal dependence arise in many scientific domains. We develop a Bayesian feature-extraction framework for spatio-temporal settings that employs Gaussian and Diffused-gamma priors to induce structured sparsity. The modeling framework specifies a general likelihood via Bregman divergence, enabling compatibility with a range of loss functions and measurement models. Posterior computation is carried out via Markov chain Monte Carlo (MCMC), and we introduce a two-stage feature-extraction procedure based on posterior samples to stabilize selection across space and time. We illustrate the method with a multi-subject electroencephalography (EEG) case study examining the relationship between chronic alcohol exposure and activity in different brain regions. We first fit binary classification models at each time point, then use false discovery rate-controlled screening and subsequent clustering in a two-stage feature-extraction pipeline to identify active brain regions. The analysis demonstrates that our proposed priors improve recovery of sparse features and enhance interpretability in the presence of spatio-temporal dependence. The framework is broadly applicable to high-dimensional, structured problems where accurate feature selection and inference are required. The code to implement the model is publicly available via GitHub.

↑