arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于机器学习的分布均衡方法用于利用真实世界数据增强临床试验

Distributional Balancing with Machine Learning for Clinical Trial Augmentation Using Real-World Data

Zern Ke, Mingshi Cui, Gemma Moran, Javier Cabrera

arXiv 2609.23524首次发表:更新:

AI 中文总结

针对临床试验中协变量分布难以均衡的问题,提出DBML方法,通过变分自编码器检测异常、重新加权匹配分布并采样对照组,在协变量均衡上优于现有匹配与加权算法。

AI 中文摘要

在临床试验中,通常采用治疗组和对照组的随机化来确保两组在平均意义上具有相似的协变量分布,从而得到无偏的因果效应估计。然而,由于招募成本、患者退出等原因,这种均衡的协变量分布在实践中难以实现。解决此问题的一种可能方案是纳入来自外部真实世界数据库的对照组患者。在本文中,我们提出了DBML方法,该方法通过匹配治疗组与潜在对照组之间的分布,而不是匹配两组之间的个体单元,从真实世界数据库中选取对照单元。DBML包含三个步骤。首先,我们使用变分自编码器检测(相对于治疗分布的)异常数据库单元。其次,我们对剩余的数据库单元重新加权,以匹配治疗组的分布。最后,我们利用这些权重为对照组采样单元。我们将所提出的方法与基于匹配的替代算法和加权算法进行比较,在协变量均衡方面取得了更优的性能。

英文摘要

In clinical trials, randomization of treatment and control groups is typically used to ensure the groups have similar covariate distributions on average, resulting in unbiased causal effect estimation. Such balanced covariate distributions are hard to achieve in practice, however, due to recruitment costs, patient dropouts, and more. One possible solution to this problem is to include control patients from external, real world databases. In this paper, we propose DBML, a method that selects control units from a real world database by matching the distribution between the treatment group and the potential control group, instead of matching units between two groups. DBML has three steps. First, we detect anomalous database units (with respect to the treatment distribution) using a variational autoencoder. Second, we re-weight the remaining database units to match the distribution of the treatments. Finally, we use these weights to sample units for our control group. The proposed method is compared to alternative matching-based algorithms and weighted algorithms, achieving superior performance in covariate balance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑