面向抗癌药物反应建模的大规模AI就绪数据
Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling
浏览论文内容
中文总结 AI 辅助
本研究扩展IMPROVE基准数据集,新增超5万种化合物,训练的DRP模型在药物盲等场景泛化能力提升,为抗癌药物反应建模提供了更丰富的社区资源。
中文摘要 AI 辅助
药物反应预测(DRP)模型是药物基因组学领域的研究热点,有望加速有效抗癌药物的筛选,但其预测性能常受限于数据集规模有限、癌症与化学空间覆盖不足的问题。此外,不一致的基准测试方法阻碍了不同模型间的可靠比较。标准化框架如用于预测肿瘤模型评估的创新方法与新数据(IMPROVE)项目,为一致的基准测试提供了统一的数据模式与评估协议,但提升模型泛化能力需要更大规模、更多样化的训练数据。本研究通过大规模整合主要来自PharmacoDB的药物基因组学数据及其他小型数据源,大幅扩展了IMPROVE基准。扩展后的资源包含数百万个药物反应测量值、更广泛的多组学覆盖范围,以及新增超过50000种化合物的化学多样性。为评估新数据集相比原始IMPROVE基准数据集的影响,使用两种数据集分别训练DRP模型,并采用通用测试集及多种评估策略(包括药物盲、癌症盲和不相交数据拆分)评估其预测性能。尽管癌症盲性能与原始基准相当,但在药物盲和不相交设置下,基于扩展数据集训练的模型表现出持续提升,表明其对先前未见过的化合物的泛化能力增强。这些结果使扩展数据集成为社区资源,为开发旨在助力新型抗癌药物发现的DRP模型提供了更丰富的基础。
英文摘要
Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer drugs. However, their predictive performance is often constrained by limited dataset scale and insufficient coverages of cancer and chemical spaces. In addition, inconsistent benchmarking practices hinder reliable comparison across models. Standardized frameworks, such as the Innovative Methodologies and New Data for Predictive Oncology Model Evaluation (IMPROVE) project, provide unified data schemas and evaluation protocols for consistent benchmarking, but improving model generalizability requires larger and more diverse training data. In this work, we substantially expand the IMPROVE benchmark through large-scale integration of pharmacogenomic data, primarily from PharmacoDB, together with additional smaller data sources. The expanded resource includes millions of drug response measurements, broader multi-omics coverage, and a major increase in chemical diversity, adding more than 50,000 compounds. To evaluate the impact of the new dataset compared to the original IMPROVE benchmark dataset, we trained DRP models using the two datasets and assess their prediction performance using a common test set and several evaluation strategies, including drug-blind, cancer-blind, and disjoint data splits. While cancer-blind performance remained comparable to the original benchmark, models trained on the expanded dataset showed consistent improvements in drug-blind and disjoint settings, indicating enhanced generalization to previously unseen compounds. These results position the expanded dataset as a community resource that provides a richer foundation for developing DRP models intended to aid in the discovery of novel anticancer drugs.
发表机构
- Argonne National Laboratory(阿贡国家实验室)
- University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。