AI 中文总结
该研究针对基因重构预训练模型未直接优化全细胞表示的问题,提出含三种适配的互补转录组视图对比预训练框架,在细胞类型注释和基因调控网络推断任务中表现出竞争力。
AI 中文摘要
单细胞转录组数据的快速增长推动了主要通过重构掩蔽表达值进行预训练的基础模型的发展。该目标鼓励这些模型学习基因依赖关系,但并未直接优化对许多下游任务至关重要的全细胞表示。为弥合这一差距,我们提出了一种对比预训练框架,通过互补转录组视图学习细胞表示。由于标准对比学习难以直接应用于单细胞预训练,我们在三个维度引入了特定适配:共表达引导的基因划分、表达感知的对比集构建,以及能力门控的对比起始。具体而言,我们首先根据每个细胞基因的共表达结构对其进行划分,构建该细胞的两个互补视图。然后,为防止模型将基因集身份作为捷径,我们通过置换表达值同时保持基因身份不变来构建难负样本。最后,我们引入能力感知控制器来确定对比目标的应用方式。在细胞类型注释和基因调控网络推断上的实验表明,在评估的协议下取得了具有竞争力的迁移性能。在六网络GRN评估中,我们的方法在所有对比变体中记录了最高的平均AUROC和AUPRC点估计,而各网络的最高得分变体有所不同。这些结果确立了互补视图对比学习作为超越基因重构的单细胞预训练的有效方向。
英文摘要
The rapid growth of single-cell transcriptomic data has enabled the development of foundation models pretrained primarily by reconstructing masked expression values. This objective encourages these models to learn gene dependencies but does not directly optimize whole-cell representations, which are essential for many downstream tasks. To bridge this gap, we propose a contrastive pretraining framework that learns cell representations through complementary transcriptomic views. Since standard contrastive learning is not readily applicable to single-cell pretraining, we introduce specific adaptations along three dimensions --- co-expression-guided gene partitioning, expression-aware contrast-set construction, and competence-gated contrastive onset. Specifically, we first construct two complementary views of each cell by partitioning its genes according to their co-expression structure. Then, to prevent the model from using gene-set identity as a shortcut, we construct hard negatives by permuting expression values while keeping gene identities unchanged. Finally, we introduce a competence-aware controller to determine how the contrastive objective is applied. Experiments on cell-type annotation and gene regulatory network inference demonstrate competitive transfer under the evaluated protocols. In the six-network GRN evaluation, our method records the highest mean AUROC and AUPRC point estimates among the compared variants, while the highest-scoring variant differs across individual networks. These results establish complementary-view contrastive learning as an effective direction for single-cell pretraining beyond gene reconstruction.
Comments9 pages