arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23900stat.ME

使用均值参数化狄利克雷模型的有限混合处理成分数据中的温和异常值与未观测值

Handling mild outliers and unobserved values in compositional datasets using finite mixtures of mean-parametrised Dirichlet models

Jason Pillay, Andriëtte Bekker, Cristina Tortora, Antonio Punzo

首次发表
浏览论文内容

中文总结 AI 辅助

针对含缺失值与异常值的成分数据,本文提出均值参数化狄利克雷模型的有限混合方法,通过定制EM算法实现聚类与异常值检测,在真实数据应用中优于传统狄利克雷混合模型。

中文摘要 AI 辅助

异质成分数据可能同时受缺失值和非典型点影响,对聚类和异常值检测均构成挑战。我们针对Huber污染模型下的不完整成分数据构建混合模型,该模型的污染直接定义在单纯形上及随机缺失的观测值上。该模型为异常值提供了原则性表示,并能在考虑污染的同时推导缺失部分的分布。我们证明,由受污染的均值参数化狄利克雷密度构建的观测对数似然最大化是凸优化问题。随后开发了定制化期望最大化(EM)算法,其E步纳入数据缺失部分分布的矩。尽管所得参数估计无闭式解,但最大化步可采用易处理的逐元素迭代更新。数值实验在不同缺失率、污染程度和样本量下验证了所提方法的性能。将其应用于美国时间使用调查数据,识别出对应工作密集型与社交型休闲日的两个可解释聚类,同时揭示了非典型时间使用成分;相比之下,传统狄利克雷混合模型识别出四个聚类,反映了异常值的影响及某一聚类的人为拆分。

英文摘要

Heterogeneous compositional data may be simultaneously affected by missing values and atypical points, posing challenges for both clustering and outlier detection. We develop a mixture model for incomplete compositional data under Huber's contamination model, with contamination defined directly on the simplex and on observations that may be missing at random. The model provides a principled representation of outliers and allows the distribution of missing parts to be derived while accounting for contamination. We establish that maximisation of the observed log-likelihood constructed from contaminated, mean-parametrised Dirichlet densities is a convex optimisation problem. We then develop a tailored expectation-maximisation. The E-step incorporates the moments from the distribution of the missing parts of the data. Although the resulting parameter estimates are not available in closed form, the maximisation step admits tractable element-wise iterative updates. Numerical experiments demonstrate the performance of the proposed approach under varying percentage of missingness and contamination, and different sample size. An application to the American Time Use Survey identifies two interpretable clusters corresponding to work-intensive and sociable recreational days, while revealing atypical time-use compositions. In contrast, a conventional Dirichlet mixture model identifies four clusters, reflecting the influence of outliers and an artificial splitting of one cluster.

↑