结合转录组数据、蛋白质结构与定位信息的亚细胞分辨率单细胞嵌入学习
Subcellularly Resolved Single-Cell Embedding Learning with Transcriptomic data, Protein Structure and Localization Information
- Institute of Image Processing and Pattern Recognition, Shanghai Jiao Tong University(上海交通大学图像处理与模式识别研究所)
- Institute of Process Engineering, Chinese Academy of Sciences(中国科学院过程工程研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出一种多模态交叉注意力框架,联合转录组、蛋白质序列与结构信息,生成亚细胞分辨率的细胞嵌入,为细胞表示学习提供了新的跨模态范式。
AI中文摘要:
现有的细胞嵌入方法主要依赖转录组或蛋白质组测量,将每个细胞表示为整体实体,从而忽略了单个分子的亚细胞定位。此外,尽管蛋白质结构信息对决定分子相互作用和功能至关重要,但这些方法很少纳入该信息。在本研究中,我们提出了一种多模态框架,用于学习亚细胞分辨率的细胞嵌入,该框架联合利用RNA表达谱、蛋白质序列表示和蛋白质结构信息。具体而言,我们采用交叉注意力架构来整合转录组、序列和结构模态,并在不同的亚细胞区室内建模它们的相互作用。所得的嵌入通过细胞的精细亚细胞组织来表示每个细胞,既捕获分子表达模式,又捕获相关蛋白质的功能特性。通过以亚细胞分辨率学习细胞表示,我们的框架保留了空间组织的生物信息,同时整合了多个分子层面的互补信号。据我们所知,这是首个在统一跨模态学习范式内联合纳入转录组信息、蛋白质序列表示和蛋白质结构知识,以生成亚细胞分辨率细胞嵌入的框架。
英文摘要:
Existing cell embedding methods predominantly rely on transcriptomic or proteomic measurements and represent each cell as a holistic entity, thereby overlooking the subcellular localization of individual molecules. Moreover, they rarely incorporate protein structural information, despite its fundamental role in determining molecular interactions and functions. In this work, we propose a multimodal framework for learning subcellularly resolved cell embeddings by jointly leveraging RNA expression profiles, protein sequence representations, and protein structural information. Specifically, we employ a cross-attention architecture to integrate transcriptomic, sequence, and structural modalities and model their interactions within distinct subcellular compartments. The resulting embeddings represent each cell through its fine-grained subcellular organization, capturing both molecular expression patterns and the functional properties of the associated proteins. By learning cell representations at subcellular resolution, our framework preserves spatially organized biological information while integrating complementary signals across multiple molecular levels. To the best of our knowledge, this is the first framework that produces subcellularly resolved cell embeddings by jointly incorporating transcriptomic information, protein sequence representations, and protein structural knowledge within a unified cross-modal learning paradigm.