发表机构
China Agricultural University(中国农业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对单细胞基础模型数据扩展收益递减问题,提出知识增强预训练方法scKITE,整合细胞注释与基因调控信息,以极少数据超越强基线模型。
AI 中文摘要
单细胞基础模型(scFMs)日益依赖于大规模转录组预训练,然而扩大预训练数据可能带来递减的收益,同时大幅增加计算成本。我们的数据扩展分析表明,纳入生物学知识(包括细胞水平的文本注释和基因水平的调控信息)提供了比单纯增加数据规模更多的扩展维度。受此观察启发,我们提出了scKITE,一个简单而有效的scFM,通过轻量级辅助解码器将细胞注释和基因调控监督整合到共享的转录组Transformer编码器中。这些解码器仅在预训练期间使用,随后被丢弃,从而产生一个富含生物学知识、适用于下游应用的通用编码器。仅使用179,067个预训练样本(即不到先前强scFM所用样本的0.5%),scKITE在多种下游任务中超越了这些模型,凸显了知识增强预训练作为生物学基础scFM的一种有前景范式的潜力。
英文摘要
Single-cell foundation models (scFMs) increasingly rely on large-scale transcriptomic pretraining, yet expanding pretraining data can yield diminishing gains while substantially increasing computational cost. Our data scaling analyses showed that incorporating biological knowledge, including cell-level text annotation and gene-level regulatory information, provided additional scaling dimension than simply increasing data size. Motivated by this observation, we present scKITE, a simple yet effective scFM that integrates cell-annotation and gene-regulatory supervision into a shared transcriptomic Transformer encoder through lightweight auxiliary decoders. These decoders are used only during pretraining and subsequently discarded, yielding a general-purpose encoder enriched with biological knowledge for downstream applications. With only 179,067 pretraining samples, i.e., less than 0.5\% of those used by previous strong scFMs, scKITE outperformed these models across diverse downstream tasks, highlighting knowledge-enhanced pretraining as a promising paradigm for biologically grounded scFMs.