基于多组学的乳腺癌预测机器学习模型基准测试
Benchmarking Machine Learning Models for Multi-Omics-Based Breast Cancer Prediction
浏览论文内容
中文总结 AI 辅助
研究针对基于多组学数据预测乳腺癌雌激素受体(ER)状态,对随机森林等经典机器学习模型进行基准测试,采用严格实验框架,结果显示RNA表达预测信号最强,多组学整合有改进,随机森林性能最佳,还选出相关基因支持模型有效性。
中文摘要 AI 辅助
雌激素受体(ER)状态是乳腺癌诊断、预后和治疗选择中的关键生物标志物。高通量测序技术的进展使多组学数据集得以生成,为计算预测任务提供补充分子信息。本研究对使用TCGA - BRCA队列中的转录组(RNA表达)、基因组(拷贝数变异;CNV)和蛋白质组(RPPA)数据预测ER状态的经典机器学习模型进行系统基准分析。采用严格实验框架确保可靠评估并防止数据泄露。在单组学和多组学设置中评估了随机森林、XGBoost等模型。结果表明RNA表达提供最强预测信号,多组学整合有适度但一致的改进。随机森林在综合多组学设置中总体性能最佳。此外,反复选择生物学相关基因支持了模型的生物学有效性。这些发现表明精心正则化的经典机器学习方法对小的高维基因组数据集仍然非常有效,且多组学整合为乳腺癌ER状态预测提供补充信息。
英文摘要
Estrogen Receptor (ER) status is a critical biomarker in breast cancer diagnosis, prognosis, and treatment selection. Recent advances in high-throughput sequencing technologies have enabled the generation of multi-omics datasets that provide complementary molecular information for computational prediction tasks. This study presents a systematic benchmarking analysis of classical machine learning models for ER status prediction using transcriptomic (RNA expression), genomic (copy number variation; CNV), and proteomic (RPPA) data from the TCGA-BRCA cohort. A rigorous experimental framework incorporating stratified train-test splitting, stratified five-fold cross-validation, class imbalance handling, and fold-specific feature selection was employed to ensure reliable evaluation and prevent data leakage. Random Forest, XGBoost, LightGBM, CatBoost, Support Vector Machines (SVM), and Logistic Regression were evaluated across single-omic and multi-omic settings. Results demonstrated that RNA expression provided the strongest predictive signal, while multi-omic integration yielded modest but consistent improvements over individual modalities. Among all evaluated approaches, Random Forest achieved the best overall performance in the integrated multi-omic setting, obtaining a balanced accuracy of 90.3\% and an ROC-AUC of 97.1\%. Furthermore, recurrent selection of biologically relevant genes, including \textit{ESR1}, \textit{PGR}, \textit{FOXA1}, and \textit{GATA3}, supported the biological validity of the learned models. These findings indicate that carefully regularized classical machine learning methods remain highly effective for small, high-dimensional genomic datasets and that multi-omic integration provides complementary information for breast cancer ER status prediction.