发表机构
St. Xavier’s College (Autonomous), Kolkata; Indian Statistical Institute(加尔各答圣泽维尔学院(自治); 印度统计研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种边缘化狄利克雷过程列聚类与尖峰- slab稀疏性的贝叶斯非参数因子模型,自适应推断因子数并合并冗余字典,实现最优收缩率,在模拟和乳腺癌数据中优于基线。
AI 中文摘要
我们提出了一种贝叶斯非参数因子模型,该模型能够推断因子数量、诱导行方向稀疏性,并通过精确聚类合并冗余字典元素。在过完备载荷矩阵的列上放置狄利克雷过程先验,并完全边缘化为精确的Pólya urn,避免了stick-breaking和辅助变量。尖峰- slab基测度允许整个因子精确为零。该模型在单一边缘化狄利克雷过程中独特地结合了精确零、列可交换性和精确合并,不同于CUSP、MGP或beta过程。我们开发了一个带有规范重标记和并行C/MPI实现的精确Gibbs采样器。我们证明了协方差矩阵的后验收缩率为$\sqrt{M s_0 \log n / n}$,在固定字典设置下获得极小极大最优率$\sqrt{s_0 \log n / n}$以及欠拟合一致性;过拟合方向是一个开放猜想。尖峰- slab是必不可少的:没有它,有效维度随$pM$缩放,导致较慢的速率。模拟表明,该方法是唯一完全自适应恢复真实秩的方法,实现了最小的协方差、载荷和信号重建误差,优于oracle基线。在van 't Veer乳腺癌数据($n=97$, $p=1213$)上,后验集中于八个可解释的程序;七个通过一致性检验,两个通过Bonferroni校正的Hallmark富集检验。
英文摘要
We propose a Bayesian nonparametric factor model that infers the number of factors, induces row-wise sparsity, and merges redundant dictionary elements via exact clustering. A Dirichlet process prior is placed on the columns of an overcomplete loading matrix and fully marginalized to an exact Pólya urn, avoiding stick-breaking and auxiliary variables. A spike-and-slab base measure allows entire factors to be exactly zero. The model uniquely combines exact zeros, exchangeability over columns, and exact merging within a single marginalized Dirichlet process, unlike CUSP, MGP, or the beta process. An exact Gibbs sampler with canonical relabeling and parallel C/MPI implementation is developed. We prove posterior contraction at rate $\sqrt{M s_0 \log n / n}$ for the covariance matrix, and in the fixed-dictionary setting obtain the minimax optimal rate $\sqrt{s_0 \log n / n}$ plus underfitting consistency; the overfitting direction is an open conjecture. The spike-and-slab is essential: without it the effective dimension scales as $pM$, yielding a slower rate. Simulations show the method is the only fully adaptive approach to recover the true rank, achieving the smallest covariance, loading, and signal-reconstruction errors, beating an oracle baseline. On van 't Veer breast cancer data ($n=97$, $p=1213$), the posterior concentrates on eight interpretable programmes; seven pass coherence and two pass Bonferroni-corrected Hallmark enrichment.
Comments"The (Factor) Matrix Reloaded" -- comments and questions welcome