arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GeneICL:面向批量转录组学的表格基础模型

GeneICL: A Tabular Foundation Model for Bulk Transcriptomics

Michael Bohl, Alexander Theus, David Wissel, Valentina Boeva

arXiv 2610.08694首次发表:更新:

发表机构

ETH Zurich; Max Planck Institute for Intelligent Systems; Swiss Institute of Bioinformatics; Université Paris Cité(苏黎世联邦理工学院; 马克斯·普朗克智能系统研究所; 瑞士生物信息学研究所; 巴黎西岱大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GeneICL是一种4.2M参数的表格基础模型,通过转录组感知的半合成预训练和循环架构,在80个临床预测任务中优于自监督转录组模型,实现高效分类、回归和生存预测。

AI 中文摘要

基因表达在生物医学中被广泛测量,但由于高维性、强特征相关性和有限的标记数据,临床结果预测仍然具有挑战性。大型自监督转录组基础模型往往无法超越简单的监督基线。表格基础模型通过上下文学习提供了一种替代方案,但通常在通用合成数据而非转录组结构上进行预训练。我们探究是否转录组感知的预训练(而非规模)才是缺失的关键要素。为此,我们引入了GeneICL,一个4.2M参数的表格基础模型,结合了从实测批量表达谱构建的半合成预训练先验与参数高效的循环架构。我们进一步通过使用Cox偏似然残差的无训练降级为回归,实现了右删失生存预测。我们在80个临床结果预测任务上评估了GeneICL,涵盖分类、回归和生存分析。表格基础模型持续优于自监督转录组模型,而GeneICL在评估的基础模型和调优基线中取得了最佳总体排名。GeneICL以最多387倍的参数减少、推理时无梯度更新以及在笔记本电脑CPU上秒级预测实现了这一性能。

英文摘要

Gene expression is widely measured in biomedicine, yet clinical outcome prediction remains challenging due to high dimensionality, strong feature correlations, and limited labeled data. Large self-supervised transcriptomic foundation models often fail to outperform simple supervised baselines. Tabular foundation models offer an alternative through in-context learning, but are typically pretrained on generic synthetic data rather than transcriptomic structure. We ask whether transcriptomics-aware pretraining, rather than scale, is the missing ingredient. Towards this end, we introduce GeneICL, a 4.2M-parameter tabular foundation model combining a semi-synthetic pretraining prior built from measured bulk expression profiles with a parameter-efficient recurrent architecture. We further enable right-censored survival prediction via a training-free reduction to regression using Cox partial-likelihood residuals. We evaluate GeneICL on 80 clinical outcome-prediction tasks spanning classification, regression, and survival. Tabular foundation models consistently outperform self-supervised transcriptomic models, while GeneICL achieves the best overall rank among evaluated foundation models and tuned baselines. GeneICL does so with up to 387$\times$ fewer parameters, no gradient updates at inference, and predictions within seconds on a laptop CPU.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑