arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06659cs.AI

CellWorld:空间转录组学基础模型中从基因水平重建到潜在细胞预测

CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models

Haiping Liu, Qian Zhao, Lijing Lin, Jingyuan Sun, Hongpeng Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出CellWorld,将空间转录组学基础模型的预测目标从基因测量转向潜在细胞表示,在多基准测试中其性能优于现有基线,仅用5%语料库预训练的大型模型也表现出色。

中文摘要 AI 辅助

本文表明,潜在空间预测性预训练可为空间转录组学基础模型提供可扩展的途径。现有空间转录组学基础模型主要重建被掩盖的基因身份或表达值,这可能会促使重现检测特定的技术变异,并限制表示的可迁移性。为避免直接重建此类变异,我们将预测目标从观测到的基因测量值转移到潜在细胞表示,并引入CellWorld,该模型从可见空间上下文和有限的部分表达提示中预测被掩盖细胞的潜在表示。我们在包含4600万个人类细胞的语料库上预训练了四个CellWorld变体,其可训练参数范围为574万至9456万。我们的可控缩放实验表明,性能随模型容量提高而提升,尤其在空间任务上,而空间迁移更多取决于充分的优化和广泛的生物源多样性,而非仅细胞数量。在四个保留的数据集上,即使是具有574万可训练参数的CellWorld-Small,在全部11个线性探测基准和全部7个微调空间基准上均优于所有基线。最值得注意的是,仅使用5%的语料库且具有广泛生物源覆盖的冻结CellWorld-Large,在全部7个空间基准上均优于所有完全微调的基线。代码可在该https URL获取。

英文摘要

This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Existing spatial transcriptomics foundation models primarily reconstruct masked gene identities or expression values, potentially encouraging the reproduction of assay-specific technical variation and limiting representation transferability. To avoid directly reconstructing such variation, we shift the prediction target from observed gene measurements to latent cell representations and introduce CellWorld, which predicts the latent representations of masked cells from visible spatial context and a limited partial-expression hint. We pretrain four CellWorld variants, spanning 5.74M to 94.56M trainable parameters, on a corpus of 46 million human cells. Our controlled scaling experiments show that performance improves with model capacity, particularly on spatial tasks, while spatial transfer depends more on sufficient optimization and broad biological source diversity than on cell count alone. Across four held-out datasets, even CellWorld-Small, with 5.74M trainable parameters, outperforms every baseline on all 11 linear-probe benchmarks and all seven fine-tuned spatial benchmarks. Most notably, a frozen CellWorld-Large pretrained on only 5\% of the corpus with broad biological source coverage outperforms every fully fine-tuned baseline across all seven spatial benchmarks. Code is available at https://github.com/UoM-HealthAI/CellWorld.

↑