再探基于METABRIC数据预测分子乳腺癌亚型
Another look at predicting molecular breast cancer subtypes from the METABRIC data
浏览论文内容
中文总结 AI 辅助
研究利用最近收缩质心和套索方法对METABRIC乳腺癌数据进行分析,评估多个模型分类性能及生存影响,发现一对多模型误分类误差最低,多项分类任务拆分为套索二项分类器对患者聚类效果良好。
中文摘要 AI 辅助
基于基因组数据将患者分类到不同聚类中,能为其疾病特异性生存轨迹提供有价值见解。本文将最近收缩质心和套索这两种监督学习方法应用于1980名患者和754个基因的METABRIC乳腺癌数据,用R包实现相关模型并划分数据为发现集和验证集,评估模型分类性能及生存影响。发现一对多模型误分类误差最低,为0.0572,且多项分类任务拆分为多个套索二项分类器对患者聚类有良好结果。
英文摘要
Classifying patients into different clusters based on genomic data can offer valuable insights into their projected disease-specific survival trajectories over time. Here we apply two supervised learning methods---Nearest Shrunken Centroids and LASSO--- to the METABRIC Breast Cancer data {metabric} of 1980 patients and 754 genes to perform this task. The {pamr} R package implements the Nearest Shrunken Centroids classifier and the { glmnet} R package is used to fit an ungrouped multinomial model, a grouped multinomial model, and a One-Versus-Rest model. Splitting our data into discovery and validation sets, we evaluate all four models' classification performance and the survival implications of their class predictions using cross validation and Kaplan-Meier curves. We find that the One-Versus-Rest model produces the lowest misclassification error of 0.0572 and the lowest median log-rank test of 0.380 statistic measuring the similarity between its Kaplan-Meier curves and the discovery set's true Kaplan-Meier curves. We show that a multinomial classification task split into several LASSO binomial classifiers offers promising results for patient clustering.