发表机构
University of Sharjah; University of Birmingham; University Medical Center Schleswig-Holstein(沙迦大学; 伯明翰大学; 石勒苏益格-荷尔斯泰因大学医学中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用XGBoost和半监督学习构建多阶段肝细胞癌基因组数据集,含770个样本、五类标签,分类准确率达96.5%。
AI 中文摘要
肝癌是一种复杂的疾病,每年在全球造成大量死亡,使得自动化肝癌分类解决方案变得紧迫。肝癌最常见的形式是肝细胞癌(HCC),占肝癌病例的90%以上。目前明显缺乏利用基因组数据的公开HCC数据集,而这对于训练人工智能(AI)模型进行自动化HCC分类是必要的。本研究提出使用XGBoost和半监督学习在三个独立的基因组生物标志物数据集上构建一个多阶段HCC数据集,利用这些数据集在半监督学习过程中的现有标签来标注所提出的数据集。该数据集总共包含770个患者样本,分为五类,代表正常组织以及HCC的不同阶段。数据集中的每个样本包含11,150个不同的基因表达水平。XGBoost模型在半监督学习过程中最终分类准确率达到96.5%。
英文摘要
Liver cancer is a complex disease responsible for a high number of deaths across the globe each year, making automated solutions for liver cancer classification urgent. The most common form of liver cancer is hepatocellular carcinoma (HCC), accounting for over 90% of liver cancer cases. There is a distinct lack of publicly available HCC datasets utilizing genomic data, which is necessary for training artificial intelligence (AI) models for automated HCC classification. This study proposes constructing a multi-stage HCC dataset using XGBoost and Semi-Supervised learning on three separate datasets of genomic biomarkers, utilizing their existing labels in the Semi-Supervised learning process to label the proposed dataset. The proposed dataset consists of 770 patient samples in total, categorized into five classes that represent normal tissue alongside different stages of HCC. Each sample in the dataset consists of 11,150 different gene expression levels. The XGBoost model demonstrated a final classification accuracy of 96.5% during the Semi-Supervised learning process.
Comments6 pages, 7 figures, 2 tables, published at the 18th International Conference Series on Developments in eSystems Engineering
Journal ref2025 18th International Conference on Development in eSystem Engineering (DeSE), Bucharest, Romania, 2025, pp. 555-560
DOI:10.1109/DeSE68208.2025.11368219