发表机构
Westlake University; Centre for Artificial Intelligence and Robotics (CAIR), Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences (HKISI-CAS); Institute of Automation, Chinese Academy of Sciences (CASIA); University of Chinese Academy of Sciences (UCAS)(西湖大学; 中国科学院香港创新研究院人工智能与机器人中心; 中国科学院自动化研究所; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
BioM-JEPA通过预测图连接基因块的聚合表示学习单细胞嵌入,在CellBench任务中实现最低扰动响应误差,微调与保留嵌入吞吐量优于scFoundation,为单细胞表示学习提供新预测单元。
AI 中文摘要
单细胞转录组是对协调生物程序的稀疏观测,然而大多数自监督模型通过重构单个基因进行学习。本文提出BioM-JEPA,这是一种联合嵌入预测架构,它转而预测由蛋白质关联和语料库衍生的共表达证据定义的图连接基因块的聚合表示。学生网络从细胞中剩余的基因推断每个目标块的表示,而缓慢更新的教师网络从完整观测的基因集提供对应的目标。在报告的提取流程下,与 token 预测、随机块和重构对照相比,块级预测产生的嵌入具有更高的有效秩,且与检测到的基因深度的关联更弱。在CellBench任务中,冻结的BioM-JEPA嵌入保留了表达、通路和邻域信息,在评估的模型中实现了最低的聚合扰动响应误差。表示诊断也与经典胰腺程序及遗传扰动间的组成关系一致。线性注意力避免构建二次方的基因-基因注意力矩阵;在批大小为8的匹配单epoch hPancreas实验中,BioM-JEPA提供了5.75倍的微调吞吐量和3.76倍的保留嵌入吞吐量,优于scFoundation。综上,这些结果支持图连接基因块作为单细胞生物学中JEPA式表示学习的有用预测单元。
英文摘要
Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence. A student network infers each target-block representation from the remaining genes in a cell, while a slowly updated teacher supplies the corresponding target from the full observed gene set. Under the reported extraction procedure, block-level prediction produced embeddings with higher effective rank and weaker association with detected-gene depth in the tested diagnostics than token-prediction, random-block and reconstruction controls. Across CellBench tasks, frozen BioM-JEPA embeddings retained expression, pathway and neighbourhood information and achieved the lowest aggregate perturbation-response error among the evaluated models. Representation diagnostics were also consistent with canonical pancreatic programmes and compositional relationships between genetic perturbations. Linear attention avoids constructing a quadratic gene-by-gene attention matrix; in a matched one-epoch hPancreas experiment at batch size 8, BioM-JEPA provided 5.75-fold higher fine-tuning throughput and 3.76-fold higher held-out embedding throughput than scFoundation. Together, these results support graph-connected gene blocks as useful prediction units for JEPA-style representation learning in single-cell biology.
Comments34 pages, 6 figures, and 13 supplementary tables (Tables S1-S13); includes Supplementary Information with detailed training and evaluation protocols. Numerical source data for all figures are provided as ancillary files; training code and the BioM-JEPA checkpoint will be released via GitHub