为单细胞生成任务扩展自回归Transformer模型
Scaling an Autoregressive Transformer for Single-Cell Generation
浏览论文内容
中文总结 AI 辅助
本研究针对单细胞基因表达向量的自监督生成任务,扩展因果Transformer模型,发现单细胞基础模型的双指数缩放定律与计算最优前沿,还探讨其微调用于扰动响应预测的可能。
中文摘要 AI 辅助
我们研究针对单细胞基因表达向量的自监督生成任务:给定某细胞类型的一组向量,目标是生成该细胞类型的额外基因表达向量。针对该任务,我们表征了生成的基因表达向量的生物学保真度以及预训练损失的缩放行为。该模型是因果Transformer搭配学习得到的量化VAE分词器,采用交叉熵损失进行训练。为评估模型,我们以某细胞类型的留存基因表达向量为条件,生成基因表达向量,将所得的基因表达向量分布与该细胞类型的真实分布进行比较。我们通过改变训练参数数量和训练数据量,研究所提架构的缩放特性。据我们所知,我们发现了首个针对单细胞基础模型的联合拟合双指数缩放定律与计算最优前沿。最后,我们讨论该预训练模型如何可针对扰动响应预测进行微调。
英文摘要
We study a self-supervised generation task for single-cell gene expression vectors: given a set of vectors from a cell type, we aim to generate additional gene expression vectors of that cell type. For this task we characterize both the biological fidelity of the generated gene expression vectors and the scaling behavior of the pretraining loss. The model is a causal transformer paired with a learned quantized VAE tokenizer, trained with a cross-entropy loss. To evaluate the model, we condition it on held-out gene expression vectors of a cell type and generate vectors of gene expression, comparing the resulting distribution over gene expression vectors to the ground truth distribution of that cell type. We study the scaling properties of the proposed architecture by varying the number of trained parameters and the amount of training data. To our knowledge, we find the first jointly-fit two-exponent scaling law and compute-optimal frontier for a single-cell foundation model. Finally, we discuss how this pretrained model could be finetuned for perturbation response prediction.