arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种带采样重要性重采样的随机EM算法,用于非线性预测变量回归中的缺失数据

A Stochastic EM Algorithm with Sampling-Importance Resampling for Missing Data in Regression with Nonlinear Predictors

Dale S. Kim

arXiv 2609.23747首次发表:更新:

发表机构

UCLA Department of Statistics & Data Science(加州大学洛杉矶分校统计与数据科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对非线性预测变量回归中的缺失数据问题,提出基于采样重要性重采样的随机EM算法(SIR-StEM),仅需似然成比例即可处理任意非线性变换,模拟显示低偏差且覆盖率接近名义水平。

AI 中文摘要

当数据缺失时,估计具有非线性预测变量变换的回归模型具有挑战性,因为非线性通常使得缺失值的条件分布难以处理。先前的方法需要特定的非线性形式,如多项式或交互项,或依赖于可能引入偏差的近似。我们提出了一种随机EM算法,该算法使用采样重要性重采样(SIR-StEM)来处理在预测变量的任意非线性变换(无论是实质性动机的还是偶然的,如样条基展开)下的缺失数据。与需要线性性或闭式条件分布的方法不同,SIR-StEM仅需评估完整数据似然至比例常数,使其适用于广泛的非线性回归模型类别。我们构建了一种利用缺失数据模式信息以提高计算效率的算法,并建立了估计量的渐近正态性。我们在两项模拟研究中展示了该方法,一项使用参数非线性变换,另一项使用基于真实行为测量的样条基展开。结果表明,SIR-StEM产生低偏差和接近名义置信区间覆盖率,优于其他常见方法。最后,我们总结了局限性并指出了未来研究的方向。

英文摘要

Estimating regression models with nonlinear predictor transformations is challenging when data are missing, because nonlinearity typically renders the conditional distribution of the missing values intractable. Previous methods require specific nonlinear forms, such as polynomials or interactions, or rely on approximations that can induce bias. We propose a stochastic EM algorithm that uses sampling-importance resampling (SIR-StEM) to handle missing data under arbitrary nonlinear transformations of predictors, either substantively motivated or incidental, such as spline basis expansions. Unlike approaches that require linearity or closed-form conditionals, SIR-StEM only requires evaluating the complete-data likelihood up to proportionality, making it applicable across a broad class of nonlinear regression models. We construct an algorithm that makes use of missing data pattern information for computational efficiency, and establish asymptotic normality of the estimator. We demonstrate the method in two simulation studies, one with parametric nonlinear transformations and another with a spline basis expansion based on real behavioral measures. Results show that SIR-StEM yields low bias and near-nominal confidence interval coverage, outperforming other common approaches. We conclude with limitations and directions for future research.

Comments23 pages, 7 figures, 1 table

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑