CoMedBench:合成医疗数据保真度与下游效用的多源基准
CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility
浏览论文内容
中文总结 AI 辅助
CoMedBench是涵盖多源合成医疗数据的可复现基准,评估多生成器在静态表格和时序ICU任务上的保真度与下游效用,实验显示合成数据可保留多数下游信号,不同生成器表现存在差异。
中文摘要 AI 辅助
获取临床数据对于开发可靠的医疗机器学习系统至关重要,但电子健康记录的直接使用受到隐私法规、机构审查、数据使用协议以及重新识别风险的限制。合成数据有望成为一种实用的替代方案:它可以保留有用的统计和临床结构,同时减少敏感患者记录的暴露。现有研究通常仅评估单个生成器、一个数据集或单一下游任务,因此难以确定合成数据何时能够支持模型开发,何时会丢失任务关键信号。我们推出CoMedBench,这是一个可复现的基准,它在通用的临床有效性框架和统一的训练与评估引擎下评估一组生成器,涵盖既定重症监护数据集上的静态表格和时序下游任务。该基准共包含来自7个公开数据源的37个数据集-任务对,其中包括20个静态表格任务和17个时序ICU时间序列任务;7个数据源分别为3个重症监护数据库(MIMIC-III、MIMIC-IV和eICU)、UCI机器学习库、CDC BRFSS糖尿病队列(2015年)、NHANES(1999-2014年)以及pycox生存数据集(GBSG和METABRIC)。该基准通过比较在真实数据和合成数据上训练、测试的模型,同时评估统计保真度和任务效用。在这些设置中,合成训练数据保留了大部分下游信号:在表格任务中,参考生成器CoMed-CTGAN的平均AUROC效用(合成数据与真实数据的性能比)为90.6%,最强生成器CoMed-TVAE的这一比例升至97.3%;时序ICU任务难度更高且对生成器更敏感,CoMed-CTGAN的AUROC保留率为81.6%,在对不平衡敏感的AUPRC指标下仅为64.0%,而CoMed-TVAE的AUROC仍保留约95%。
英文摘要
Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constrained by privacy regulation, institutional review, data-use agreements, and the risk of re-identification. Synthetic data promises a practical alternative: it can preserve useful statistical and clinical structure while reducing exposure of sensitive patient records. Prior studies often evaluate a single generator, one dataset, or a narrow downstream task, making it difficult to know when synthetic data can support model development and when it fails to preserve task-critical signal. We introduce CoMedBench, a reproducible benchmark that evaluates a family of generators under a common clinical-validity framework and one shared training and evaluation engine, spanning static tabular and temporal downstream tasks on established critical-care datasets. In total the benchmark spans 37 dataset-task pairs across two modalities consists of 20 static tabular and 17 temporal ICU time-series-drawn from seven public data sources: three intensive-care databases (MIMIC-III, MIMIC-IV, and eICU) together with the UCI Machine Learning Repository, the CDC BRFSS diabetes cohort (2015), NHANES (1999-2014), and the pycox survival datasets (GBSG and METABRIC). The benchmark evaluates both statistical fidelity and task utility by comparing models trained and tested across real and synthetic data. In these settings, synthetic training data preserves most of the downstream signal: on tabular tasks the reference generator CoMed-CTGAN retains a mean AUROC utility (the synthetic-to-real performance ratio) of 90.6%, rising to 97.3% for the strongest generator, CoMed-TVAE. Temporal ICU tasks are harder and more generator-sensitive: CoMed-CTGAN retains 81.6% (AUROC) and only 64.0% under the imbalance-sensitive AUPRC, whereas CoMed-TVAE still retains ~95% (AUROC).