arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

处理狄利克雷混合模型中的缺失与删失问题

Handling Missingness and Censoring in Dirichlet Mixture Models

Jason Pillay, Andriette Bekker, Cristina Tortora, Antonio Punzo

arXiv 2607.27403首次发表:更新:

AI 中文总结

本文提出一种保留单纯形结构的狄利克雷混合模型似然方法,通过EM类算法同时完成参数估计与插补,经仿真及捕虏体、PM₂.₅真实数据验证,可有效处理成分数据的缺失与删失问题。

AI 中文摘要

不完整成分数据分析面临一个根本性局限:基于似然的成分模型方法通常要求成分完全观测,难以直接处理单纯形上的缺失或删失比例。因此,分析人员常丢弃部分观测的成分或将数据转换为无约束空间,可能牺牲可解释性与一致性。本文提出一种不离开单纯形的不完整成分数据似然方法,具体开发了一种期望最大化(EM)类算法,用于在存在缺失和删失成分时拟合有限狄利克雷分布混合模型。该方法同时执行参数估计与基于模型的插补,保留原始变量的成分结构与可解释性。仿真实验在日益复杂的粗化机制下评估了所提估计量与插补的性能,特别关注聚类表现与模型选择结果。结果显示,即便观测不完整,聚类表现仍具优势,且与当前案例删除的替代方法相比,模型选择指标识别正确聚类数的概率更高。通过两个具有不同粗化模式的真实数据集说明该方法的实用价值:对捕虏体数据集的分析确定了四分量狄利克雷混合模型,揭示了岩石类型与物种形成方法的可解释特征;对空气质量系统中PM₂.₅物种形成数据(包含左删失与随机缺失值)的应用,支持了四分量混合模型,该模型刻画了美国各地颗粒物的成分部分。

英文摘要

Incomplete compositional data analysis faces a fundamental limitation: likelihood-based methods for compositional models generally require fully observed compositions, making it difficult to accommodate missing or censored proportions directly on the simplex. Consequently, analysts often discard partially observed compositions or transform the data into unconstrained spaces, potentially sacrificing interpretability and coherence. This paper proposes a likelihood-based method for incomplete compositional data without leaving the simplex. Specifically, we develop an Expectation-Maximisation (EM) type algorithm for fitting finite mixtures of Dirichlet distributions in the presence of missing and censored components. The proposed approach performs parameter estimation and model-based imputation simultaneously while preserving the compositional structure and interpretability of the original variables. A simulation experiment evaluates the performance of the proposed estimators and imputations under increasingly complex coarsening mechanisms. Particular attention is paid to clustering performance, and model selection outcomes. The results showed beneficial clustering performance despite observations being incomplete, and a higher probability of model selection metrics identifying the correct number of clusters compared to current alternative of case-deletion. The practical utility of the method is illustrated using two real datasets with distinct coarsened patterns. Analysis of the xenolith dataset identifies a four-component Dirichlet mixture that reveals interpretable profiles of rock types and speciation methods. Application to PM$_{2.5}$ speciation data from the Air Quality System, containing both left-censored and missing-at-random values, supports a four-component mixture model that characterises compositional parts of particulate matter across the United States.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑