AI 中文总结
该研究提出两种聚类指导的IPW策略,通过模拟和乳腺癌患者数据验证,可改善倾向得分设定错误的稳健性,其性能取决于亚组结构、样本量和推断优先级。
AI 中文摘要
逆概率加权(IPW)被广泛用于估计观察性研究中的因果效应,但该方法依赖于倾向得分的充分设定。本文比较了三种处理治疗分配异质性的策略:标准IPW、结合聚类特定倾向得分模型的聚类增强IPW,以及将估计的聚类成员作为协变量纳入的全局倾向得分模型。通过在存在和不存在潜在聚类结构的情况下进行模拟,并在倾向得分模型的协变量设定正确和缺失的条件下,针对样本量为100至500的情况,评估了偏差、均方误差(MSE)和置信区间覆盖率。与标准IPW相比,两种聚类指导策略均降低了因协变量设定缺失导致的偏差和MSE,但二者均未具有绝对优势:当存在潜在聚类结构时,聚类增强IPW实现了更低的MSE;而全局模型在较小样本量下通常提供更低的偏差和更好的覆盖率。本文还将这些方法应用于966名接受卡铂治疗的乳腺癌患者,使用广义倾向得分估计治疗周期与超敏反应风险之间的剂量反应关系。标准分析和聚类分析产生了相似的汇总估计,而聚类分析额外提供了亚组特定估计和诊断特征。总体而言,聚类指导策略可提高对倾向得分设定错误的稳健性,其相对性能取决于亚组结构、样本量和推断优先级。
英文摘要
Inverse probability weighting (IPW) is widely used to estimate causal effects in observational studies but depends on adequate propensity-score specification. We compare three strategies for addressing treatment assignment heterogeneity: standard IPW, clustering augmented IPW with cluster specific propensity score models, and a global propensity score model including estimated cluster membership as a covariate. Through simulations with and without latent cluster structure and under correctly specified and omitted covariate propensity score models, we evaluate bias, mean squared error (MSE), and confidence interval coverage across sample sizes of 100 to 500. Both cluster informed strategies reduced bias and MSE from omitted covariate misspecification relative to standard IPW, but neither uniformly dominated: clustering augmented IPW achieved lower MSE when latent cluster structure was present, whereas the global model generally provided lower bias and better coverage at smaller sample sizes. We also apply the methods to 966 breast cancer patients treated with carboplatin, using generalized propensity scores to estimate the dose response relationship between treatment cycles and hypersensitivity reaction risk. Standard and clustered analyses produced similar pooled estimates, while clustering additionally provided subgroup specific estimates and diagnostic profiles. Overall, cluster informed strategies may improve robustness to propensity score misspecification, with relative performance depending on subgroup structure, sample size, and inferential priorities.