arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20000stat.MEstat.CO

大规模数据集变量选择的一种贝叶斯双向分割框架

A Bayesian Bi-Directional Splitting Framework for Variable Selection in Large Datasets

Aaron Coats, Vinny Davies, Mayetri Gupta

AI总结:

提出一种贝叶斯双向分割框架,通过分而治之并行处理大规模数据,在保持变量选择性能的同时提升计算效率,并应用于H3N2流感数据。

AI中文摘要:

现代表格数据集在样本数量和协变量数量上正变得越来越大,这给贝叶斯变量选择带来了显著的计算负担。尽管关于将贝叶斯推断扩展到大量观测或高维协变量空间的文献众多,但在实际场景中同时处理两个维度都很大的可扩展贝叶斯变量选择研究相对较少。本文提出了一种新颖的贝叶斯变量选择框架,用于高效分析具有大量行和列的数据。该框架采用分而治之的方法,将数据沿两个方向分割成批次,并并行独立分析每个批次,之后通过两阶段共识程序将结果合并。实验证明了计算效率的提升,同时保持了强大的变量选择性能,在噪声高维设置中成功识别出相关信号,尽管数据分割导致信息损失。还提供了实用指南,包括调整参数选择和有效数据划分策略的建议,并随后应用于H3N2流感数据集。

英文摘要:

Modern tabular datasets are becoming increasingly large, both in the number of samples and covariates, posing significant challenges for Bayesian variable selection due to the resulting computational burden. While there is extensive literature on scaling Bayesian inference to large numbers of observations or high-dimensional covariate spaces, comparatively little work addresses scalable Bayesian variable selection when both dimensions are large simultaneously in a practical setting. This paper presents a novel Bayesian variable selection framework for efficiently analysing data with a large number of both rows and columns. The proposed framework operates via a divide-and-conquer approach, splitting data into batches along both directions and analysing each batch independently in parallel, after which the results are combined together in a two-phase consensus procedure. Experiments demonstrate the computational gains while retaining strong variable selection performance, successfully identifying relevant signals in noisy, high-dimensional settings despite the loss of information induced by data splitting. Practical guidelines are also provided, including recommendations for tuning parameter choices and effective data partitioning strategies, followed by a real application to an H3N2 influenza dataset.

↑