发表机构
The University of Queensland; Guangdong University of Finance and Economics(昆士兰大学; 广东财经大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对分量分布兼具重尾与方向不对称性的高维聚类问题,提出深度偏t混合模型,经模拟与真实数据验证,其在偏度较强、重尾及小样本等场景下聚类性能优于相关基准模型。
AI 中文摘要
当分量分布同时具有重尾和方向不对称性时,高维聚类具有挑战性。我们提出深度偏t混合模型(DStMM),这是一种基于广义双曲偏t正态均值-方差表示的分层因子分析混合模型。共享逆伽马混合变量沿每个完整潜在路径传播,可联合建模重尾和方向不对称性,同时保留条件高斯性,因此每个完整路径都具有精确的GHST边际表示。我们正式推导其向对称深度t、高斯深度混合和单层GHST因子分析模型的归约,讨论局部非可识别性和实现层面的参数计数约定,并推导用于估计的条件广义逆高斯分布。估计通过随机/蒙特卡洛EM算法进行,采用基于实现的显式参数计数用于BIC架构比较。模拟研究表明,当不存在偏度时,DStMM的性能与对称鲁棒模型相似,但随着方向不对称性增强,尤其是在重尾情况下,其性能提升会增大;在较小样本和不等混合比例下,同样的定性行为依然存在。两个真实数据应用提供了补充证据:在UCI手写数字基准上,DStMM在通用深度架构下聚类性能最强;在气体传感器阵列漂移数据上,DStMM优于深度高斯和深度t替代模型,且在已实现的BIC准则下选择了非平凡的第二混合层。这些结果共同支持了通过深度潜在混合传播偏度和重尾变异,同时保留精确路径级似然的价值。
英文摘要
High-dimensional clustering is challenging when component distributions are both heavy-tailed and directionally asymmetric. We propose a deep skew-$t$ mixture model (DStMM), a hierarchical factor-analytic mixture based on the generalised-hyperbolic skew-$t$ normal mean--variance representation. A shared inverse-gamma mixing variable is propagated along each complete latent pathway, allowing heavy tails and directional asymmetry to be modelled jointly while preserving conditional Gaussianity. Each complete pathway therefore admits an exact GHST marginal representation. We formalise the reductions to symmetric deep $t$, Gaussian deep-mixture, and single-layer GHST factor-analytic models, discuss local non-identifiability and the implementation-level parameter-counting convention, and derive the conditional generalised inverse Gaussian law used for estimation. Estimation is carried out by a stochastic/Monte Carlo EM algorithm, with an explicit implementation-based parameter count for BIC architecture comparison. Simulation studies show that DStMM performs similarly to the symmetric robust model when skewness is absent but provides increasing gains as directional asymmetry becomes stronger, particularly under heavier tails; the same qualitative behaviour persists under smaller samples and unequal mixture proportions. Two real-data applications provide complementary evidence. On the UCI handwritten-digit benchmark, DStMM gives the strongest clustering performance under a common deep architecture, while on the Gas Sensor Array Drift data, DStMM improves on both deep Gaussian and deep $t$ alternatives and, under the implemented BIC criterion, selects a non-trivial second mixture layer. Together, these results support the value of propagating skewness and heavy-tail variation through a deep latent mixture while retaining an exact pathway-level likelihood.