arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

存在随机缺失响应与协变量时高斯混合模型的新视角

A New Look at Gaussian Mixtures in the Presence of Missing-at-Random Responses and Covariates

Hung Tong, Antonio Punzo, Cristina Tortora

arXiv 2608.03757首次发表:更新:

AI 中文总结

本文针对响应与协变量均含随机缺失值的情况,提出基于EM算法的多元线性回归混合模型,可实现不完整数据下的软聚类,其有效性经模拟与Automobile数据集验证。

AI 中文摘要

缺失值是统计建模中常见的挑战,因此恰当处理缺失值是重要的研究方向。在产生缺失值的各类机制中,最常见的是随机缺失(MAR)机制,即缺失概率仅取决于观测数据,与未观测数据无关。本文在最大似然(ML)框架下,解决了响应空间和协变量空间均存在MAR值时,含多个随机协变量的多元线性回归模型的估计问题。所提方法通过多元高斯分布的条件-边缘分解对响应与协变量的联合分布进行建模;当变量可自然划分为响应和协变量时,该公式可解释为多元正态分布的重参数化。参数估计采用期望最大化(EM)算法,该算法在保留响应与协变量不同角色的同时,便于缺失值的插补。我们将该框架扩展到基于模型的聚类场景,考虑含多个随机协变量的多元线性回归混合模型;该扩展支持不完整数据下的软聚类,且可容纳多元响应和协变量中的MAR值,因此是当前文献中回归数据最通用的基于模型的聚类解决方案之一。通过模拟研究验证了该方法的有效性,并利用含缺失值的Automobile数据集说明了所提重参数化的优势。

英文摘要

Missing values present a common challenge in statistical modeling, so handling them properly is an important research direction. Among the various mechanisms that can generate missing values, the most common is the missing-at-random (MAR) mechanism, in which the probability of missingness depends only on observed data and not on unobserved data. This paper addresses the problem of estimating a multivariate linear regression model with multiple random covariates in the presence of MAR values in both the response and covariate spaces using a maximum likelihood (ML) framework. The proposed methodology models the joint distribution of responses and covariates through a conditional-marginal factorization of a multivariate Gaussian distribution. This formulation can be interpreted as a reparameterization of the multivariate normal distribution when the variables can be naturally partitioned into responses and covariates. Parameter estimation is performed using the expectation-maximization (EM) algorithm, which facilitates the imputation of missing values while preserving the distinct roles of responses and covariates. We extend this framework to the model-based clustering setting by considering a mixture of multivariate linear regressions with multiple random covariates. This extension enables soft clustering under incomplete data and accommodates MAR values in both the multivariate responses and covariates. Hence, it represents one of the most general model-based clustering solutions for regression data currently available in the literature. The effectiveness of the methodology is demonstrated through a simulation study, and the advantages of the proposed reparameterization are illustrated using the Automobile dataset, which contains missing values.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑