用于高维混合模型的贝叶斯蒸馏聚类
Bayesian Distilled Clustering for High-Dimensional Mixture Models
浏览论文内容
中文总结 AI 辅助
本文针对高维混合模型提出贝叶斯蒸馏聚类框架,通过贝叶斯变量选择确定协变量子空间并在其上聚类,以识别分布与行为存在差异的潜在亚群,解决高维聚类的噪声干扰问题。
中文摘要 AI 辅助
潜在亚群分析是基因组学、精准医学和社会科学等领域的核心任务,其目标是识别具有不同协变量结构或响应行为的异质群体。混合模型为该任务提供了自然的概率框架,将数据生成分布表示为带有未观测标签的亚群特定分布的加权组合。在高维场景下,这类分析面临重大挑战:通常仅有一小部分协变量驱动有意义的亚群分离,其余变量可能引入噪声或冗余;标准聚类方法通常将所有维度视为同等重要,但在高维空间中,无关坐标会扭曲距离并掩盖定义潜在类别的低维结构。本文提出一种用于高维混合模型的贝叶斯蒸馏聚类框架,认为聚类应在统计上合理的子空间而非完整环境空间中进行。该方法利用贝叶斯变量选择模型估计后验包含概率,量化每个协变量对亚群分离或响应行为有贡献的证据;随后通过控制预期错误发现比例确定“蒸馏”后的协变量集,在该降维子空间上执行聚类,再通过条件独立性诊断检验所选变量间的亚群特定依赖关系。关键在于,该框架是基于模型的:蒸馏步骤直接与混合结构和响应模型绑定,而非通用降维准则,这确保所得子空间始终契合科学目标,即识别在分布结构和行为上均存在差异的潜在亚群。
英文摘要
Latent subgroup analysis is central to fields such as genomics, precision medicine, and social science, where the goal is to identify heterogeneous populations with distinct covariate structures or response behaviors. Mixture models provide a natural probabilistic framework for this task, representing the data-generating distribution as a weighted combination of subgroup-specific laws with unobserved labels. In high-dimensional regimes, these analyses face significant challenges. Often, only a small subset of covariates drives meaningful subgroup separation; the remaining variables may introduce noise or redundancy. Standard clustering methods typically treat all dimensions as equal, but in high-dimensional spaces, irrelevant coordinates can distort distances and obscure the low-dimensional structures defining latent classes. This paper introduces a Bayesian distilled clustering framework for high-dimensional mixture models. We propose that clustering should occur within a statistically justified subspace rather than the full ambient space. Our method utilizes a Bayesian variable selection model to estimate posterior inclusion probabilities, quantifying the evidence that each covariate contributes to subgroup separation or response behavior. A "distilled" covariate set is then identified by controlling the expected false-discovery proportion. Clustering is performed on this reduced subspace, followed by conditional independence diagnostics to examine subgroup-specific dependencies among selected variables. Critically, this framework is model-based: the distillation step is tied directly to the mixture structure and response model rather than a generic dimension-reduction criterion. This ensures the resulting subspace remains aligned with the scientific objective: identifying latent subgroups that differ in both distributional structure and behavior.