arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CellMSA:用于单细胞表示学习的上下文建模

CellMSA: Context Modeling for Single-Cell Representation Learning

Suyuan Zhao, Minghao Liu, Yizhen Luo, Zaiqing Nie

arXiv 2609.38908首次发表:更新:

发表机构

Institute for AI Industry Research (AIR), Tsinghua University; Department of Computer Science and Technology, Tsinghua University; Tsinghua Institute of Multidisciplinary Biomedical Research (TIMBR), Tsinghua University; National Institute of Biological Sciences (NIBS); PharMolix Inc.(清华大学人工智能产业研究院(AIR); 清华大学计算机科学与技术系; 清华大学多学科生物医学研究院(TIMBR); 国家生物科学研究所(NIBS); PharMolix公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CellMSA利用多序列比对启发,通过跨批次和细胞类型检索上下文,构建基因对表示,提升单细胞表示学习性能,在多个基准上优于现有方法。

AI 中文摘要

单细胞转录组学能够以前所未有的分辨率描绘细胞状态,但其高维度、稀疏性以及技术批次效应给表示学习带来了重大挑战。现有的单细胞基础模型通常独立编码每个细胞,或仅对同一批次内的细胞进行建模以进行去噪,从而未能充分利用跨批次和细胞类型的丰富关系信息来建模基因表达模式。我们认为,单细胞模型可以从更具信息量的细胞上下文建模中受益。通过比较细胞间的一致性和变异性,模型可以捕获与细胞状态相关的细粒度基因-基因依赖关系,这对于学习高质量表示至关重要。受蛋白质建模中多序列比对(MSA)上下文使用的启发,我们提出了CellMSA,一个将MSA启发的归纳偏置引入转录组建模的单细胞表示学习框架。对于每个目标细胞,CellMSA从不同批次和生物学相关的细胞类型中检索相关细胞作为上下文,并将跨细胞模式总结为上下文相关的基因对表示。然后,该表示被注入到一个对感知的目标细胞编码器中,以进行细粒度的表示学习。我们在一个包含约1.09亿个细胞观测数据(其中包括6560万个初级观测数据)的大规模人类单细胞语料库上预训练了CellMSA。实验表明,我们的框架在多个基准上持续优于现有方法。代码可在以下仓库获取:此https URL。

英文摘要

Single-cell transcriptomics enables profiling of cellular states at unprecedented resolution, but its high dimensionality, sparsity, and technical batch effects pose significant challenges for representation learning. Existing single-cell foundation models typically encode each cell independently or only model cells from the same batch for denoising, thereby underutilizing the rich relational information across batches and cell types to model gene expression patterns. We argue that single-cell models can benefit from more informative cell-context modeling. By comparing consistency and variation across cells, models can capture fine-grained gene-gene dependencies associated with cell states, which are essential for learning high-quality representations. Inspired by the use of multiple sequence alignment (MSA) context in protein modeling, we propose CellMSA, a single-cell representation learning framework that introduces an MSA-inspired inductive bias into transcriptomic modeling. For each target cell, CellMSA retrieves relevant cells from different batches and biologically related cell types as context, and summarizes cross-cell patterns into a context-dependent gene-pair representation. This representation is then injected into a pair-aware target-cell encoder for fine-grained representation learning. We pretrain CellMSA on a large-scale human single-cell corpus of approximately 109 million cell observations, including 65.6 million primary observations. Experiments show that our framework consistently outperforms existing methods across multiple benchmarks. Code is available at the following repository: https://github.com/PharMolix/CellMSA.

CommentsAccepted by NeurIPS 2026, code released

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑