AI 中文总结
该研究开发了代谢组学专用大语言模型MetaboLLM及配套的MetaboLLM-GIN,经多阶段适配后,在相关任务及两个临床预测场景中均取得优于对照模型的性能,证明领域专用语言模型可构建可预测且可解释的代谢物图表征。
AI 中文摘要
代谢组学知识分散在异构资源中,难以转化为可预测的表征。我们开发了MetaboLLM,这是一种通过持续预训练、监督微调及结构化检索适配的代谢组学专用大语言模型,同时开发了MetaboLLM-GIN,该模型利用图同构网络将生成的生化描述转化为代谢物图,用于患者水平预测。在四个主干模型系列上,MetaboLLM在代谢组学知识、关系及描述任务上的表现优于对应的基础模型和医学适配模型,且可迁移至外部公共基准。MetaboLLM-GIN在冠状动脉旁路移植术后应激性高血糖预测中获得最高AUC(0.8616),在绝经后激素治疗方案分类中获得最高AUC(0.8123),优于传统模型、替代图构建方式,以及未适配或未使用检索的LLM配置生成的图。模型解释在两个应用中均产生了具有生物学意义的发现。这些结果表明,领域专用语言模型可将异构生化知识组织为可预测且可解释的代谢物图表征。
英文摘要
Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. We developed MetaboLLM, a metabolomics-specialized large language model adapted through continual pretraining, supervised fine-tuning, and structured retrieval, together with MetaboLLM-GIN, which converts generated biochemical descriptions into metabolite graphs for patient-level prediction using a graph isomorphism network. Across four backbone families, MetaboLLM outperformed corresponding base and medically adapted models on metabolomics knowledge, relational, and description tasks, and transferred to an external public benchmark. MetaboLLM-GIN achieved the highest AUC for stress hyperglycemia prediction after coronary artery bypass grafting (0.8616) and postmenopausal hormone-regimen classification (0.8123), outperforming conventional models, alternative graph constructions, and graphs generated from unadapted or non-retrieval LLM configurations. Model interpretation further produced biologically meaningful findings in both applications. These results show that domain-specialized language models can organize heterogeneous biochemical knowledge into predictive and interpretable metabolite graph representations.
Comments60 pages, 3 figures, 16 tables; includes Supplementary Information