发表机构
University of Washington; University of Florida; Wake Forest University; Kaiser Permanente Washington Health Research Institute; Fred Hutchinson Cancer Center(华盛顿大学; 佛罗里达大学; 维克森林大学; 凯撒永久华盛顿健康研究所; 弗雷德·哈钦森癌症中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究开发了R包rw,提出适用于预测均值匹配的近似RW方差程序,经模拟验证其在参数和PMM插补下可使多重插补的方差估计更接近名义覆盖率。
AI 中文摘要
多重插补(MI)广泛应用于医学研究中处理缺失数据,其方差通常使用鲁宾规则(RR)进行估计。RR应用简便,但在插补程序与分析程序不兼容时,其方差估计值可能过小或过大,导致置信区间校准错误。Robins和Wang(RW)开发了一种估计方程方差估计量,在其参数插补框架的条件下不要求兼容性。然而,RW的实际应用因缺乏易用软件以及现代插补工作流程所需的额外方差计算而受限。我们开发了R包rw,为通过mice拟合的正态和逻辑条件插补模型构建RW方差分量,并提出了适用于预测均值匹配(PMM)的近似RW方差程序。该PMM程序保留原始插补值和选定的供体。对于方差计算,我们定义可微的供体概率,并使用期望分析得分近似分析得分与插补得分之间的交叉矩;按记录的供体对分析得分求和,以考虑供体重用。我们基于原始RW设置的基准模拟,以及受电子健康记录数据启发的验证研究模拟,评估所提方法。在部分设置中,RR的覆盖率较名义95%水平低约23个百分点,在其他设置中则高于名义水平。所提RW估计量在参数插补和PMM插补下,于较大样本中通常产生更接近名义水平的覆盖率,同时保留与RR相同的完整数据点估计值。
英文摘要
Multiple imputation (MI) is widely used for handling missing data in medical research, with variance commonly estimated using Rubin's rules (RR). RR is simple to apply, but under uncongeniality between the imputation and analysis procedures, its variance estimate can be either too small or too large, leading to miscalibrated confidence intervals. Robins and Wang (RW) developed an estimating-equation variance estimator that does not require congeniality under the conditions of their parametric imputation framework. Practical use of RW, however, has been limited by the lack of accessible software and by the additional variance calculations required for modern imputation workflows. We develop an R package, rw, construct RW variance components for normal and logistic conditional imputation models fitted through mice, and propose an approximate RW variance procedure for predictive mean matching (PMM). The PMM procedure preserves the original imputations and selected donors. For the variance calculation, we define differentiable donor probabilities and use expected analysis scores to approximate the cross-moment between analysis and imputation scores. Analysis scores are summed by recorded donor to account for donor reuse. We evaluate the proposed methods in benchmark simulations based on the original RW setting and in validation-study simulations motivated by electronic health record data. RR coverage was up to approximately 23 percentage points below the nominal 95% level in some settings and above nominal in others. The proposed RW estimators generally produced coverage closer to the nominal level in larger samples under both parametric and PMM imputation, while retaining the same completed-data point estimates as RR.
Comments33 pages, 5 figures, 6 tables