arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18914math.STstat.COstat.MEstat.TH

基于复合散度的单元式与个案式污染下的稳健多元估计方法

A Composite Divergence Approach to Robust Multivariate Estimation under Cellwise and Casewise Contamination

Abhik Ghosh, Claudio Agostinelli, Ayanendranath Basu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出复合密度幂散度(CDPD),构建最小CDPD估计量(MCDPDE),结合复合似然的计算可扩展性与密度幂散度的稳健性,可同时抵御单元式与个案式污染,且已通过R包mvdpd实现,模拟与实证显示其稳健性优于极大复合似然估计且效率相当。

中文摘要 AI 辅助

复合似然(CL)方法通过将联合似然替换为低维边际或条件分量的乘积,为复杂多元模型提供了比全似然推断计算效率更高的替代方案。然而,与极大似然估计(MLE)类似,极大复合似然估计(MCLE)对数据污染高度敏感。另一方面,基于稳健散度的方法,如最小密度幂散度(DPD)估计量,需要完整的联合密度,因此难以扩展到复杂多元模型。我们引入复合DPD(CDPD),这是一种完全由定义复合似然的低维分量密度构建的真正统计散度,结合了复合似然的计算可扩展性与DPD的稳健性。所得的最小CDPD估计量(MCDPDE)对MCLE进行了稳健化处理,无需在整个多元样本空间上进行积分。我们仅在分量模型的正则条件下,而非要求完整联合分布的正确设定,证明了MCDPDE的一致性、渐近正态性及影响函数。我们表明,对于其调优参数的每个正值,MCDPDE均具有定性稳健性,而其极限形式为MCLE。由于其分量可在成对或单元层面选择,该框架可同时防范个案式与单元式污染。该方法直接基于分量密度而非椭圆距离结构运行,将稳健推断从现有大多数单元式稳健方法所局限的椭圆模型扩展至其他模型。我们开发了计算算法,并在配套的R包mvdpd中实现。模拟研究与实际数据应用表明,MCDPDE相较于MCLE获得了显著的稳健性提升,同时在假设模型下保持了有竞争力的效率。

英文摘要

Composite likelihood (CL) methods provide a computationally efficient alternative to full likelihood inference for complex multivariate models by replacing the joint likelihood with a product of lower-dimensional marginal or conditional components. Like the MLE, however, the maximum CL estimator (MCLE) is highly sensitive to data contamination. On the other hand, robust divergence-based procedures such as the minimum density power divergence (DPD) estimator require the full joint density and so scale poorly to complex multivariate models. We introduce the composite DPD (CDPD), a genuine statistical divergence built entirely from the low-dimensional component densities defining a CL, combining the computational scalability of CL with the robustness of the DPD. The resulting minimum CDPD estimator (MCDPDE) robustifies the MCLE without requiring integration over the full multivariate sample space. We establish consistency, asymptotic normality, and the influence function of the MCDPDE under regularity conditions on the component models alone, without requiring correct specification of the full joint distribution. We show that it is qualitatively robust for every positive value of its tuning parameter, unlike the MCLE recovered as the limit. Because its components can be chosen at the pairwise or cell level, the framework guards simultaneously against casewise and cellwise contamination. Operating directly on component densities rather than elliptical distance structures, it extends robust inference beyond the elliptical models to which most existing cellwise-robust procedures are confined. We develop computational algorithms implemented in the accompanying R package mvdpd. Simulation studies and real-data applications show that the MCDPDE achieves substantial robustness gains over the MCLE while retaining competitive efficiency under the assumed model.

↑