AI 中文总结
本文提出统一框架,扩展CG-CWM模型以处理多元回归中的异质性、缺失与污染问题,通过ECM算法完成估计,经数值研究与实际数据验证性能。
AI 中文摘要
缺失值、异常观测值以及潜在组间异质性是回归数据复杂性的常见来源。污染高斯簇加权模型(CG-CWM)为基于模型的聚类中处理异常观测值(包括异常值和杠杆点)提供了自然框架。我们将CG-CWM扩展至响应变量和协变量空间均存在随机缺失(MAR)值的数据。所提模型在回归分析中提供聚类,同时区分典型观测值、异常值以及良好和不良杠杆点。通过将协变量视为随机变量,该模型保留分配依赖性,使其可直接助力聚类形成。最大似然估计通过期望条件最大化(ECM)算法执行,该算法考虑四类不完整信息:缺失的响应变量与协变量、未知成分成员身份以及潜在污染指标。以这些指标为条件,响应变量与协变量的联合分布为多元高斯分布,可得出缺失值的闭式条件分布,并将缺失不确定性直接纳入参数更新。因此,缺失值在模型拟合过程中处理,而非通过初步插补。该框架提供聚类、分簇回归、MAR值的基于模型处理以及异常观测值检测。其性能通过不同污染水平和缺失模式下的数值研究及实际数据应用进行评估。
英文摘要
Missing values, atypical observations, and heterogeneity across latent groups are common sources of complexity in regression data. The contaminated Gaussian cluster-weighted model (CG-CWM) provides a natural framework for handling atypical observations, including outliers and leverage points, in model-based clustering. We extend the CG-CWM to data with missing-at-random (MAR) values in both the response and covariate spaces. The proposed model provides clustering in regression analysis while distinguishing typical observations, outliers, and good and bad leverage points. By treating covariates as random, the model preserves assignment dependence, allowing them to contribute directly to cluster formation. Maximum likelihood estimation is performed through an expectation-conditional maximization (ECM) algorithm that accounts for four sources of incomplete information: missing responses and covariates, unknown component memberships, and latent contamination indicators. Conditional on these indicators, the joint distribution of responses and covariates is multivariate Gaussian, yielding closed-form conditional distributions for missing values and incorporating missingness uncertainty directly into parameter updates. Thus, missing values are handled within model fitting rather than by preliminary imputation. The framework provides clustering, clusterwise regression, model-based treatment of MAR values, and detection of atypical observations. Performance is assessed through numerical studies under varying levels of contamination and missingness patterns, and a real data application.