面向可扩展的上下文感知单细胞空间转录组学组织学图像预测
Towards Scalable Context-Aware Single-Cell Spatial Transcriptomics Prediction from Histology Images
浏览论文内容
中文总结 AI 辅助
CELLO框架通过单次病理基础模型前向传播和距离衰减交叉注意力,实现从H&E图像高效预测单细胞基因表达,较DeepSpot2Cell提速14倍并提升准确性。
中文摘要 AI 辅助
从H&E染色的组织学图像预测基因表达为昂贵的空间转录组学提供了一种可扩展的替代方案,然而现有大多数方法在点(spot)级别上运行,其中来自多个细胞的信号被聚合,关键的细胞异质性被掩盖。将这一范式扩展到单细胞分辨率并非易事。直接应用病理学基础模型面临尺度不匹配的问题:其斑块级表示混合了多个细胞,而逐细胞裁剪或调整大小会扭曲形态并移除局部上下文。相反,缺乏强预训练视觉编码器的基于分割的模型往往缺乏准确分子预测所需的形态表示能力,并继承了不完美细胞边界掩码的错误。在此,我们提出CELLO,一个高效的端到端框架,对每张图像执行一次病理学基础模型前向传播,并使用网格采样同时为所有细胞提取位置特定特征。我们进一步引入一个距离衰减交叉注意力模块,利用空间偏置的局部形态上下文细化每个细胞的表示。使用来自HEST-1k的52对公开Xenium-H&E数据,涵盖12个器官和约1000万个细胞,CELLO在评估基线上提高了平均预测准确性,同时将平均全切片推理时间相比DeepSpot2Cell减少了14.0倍(平均加速,不包括上游细胞分割)。我们的工作为从H&E图像进行单细胞基因表达预测建立了可扩展的基础。
英文摘要
Predicting gene expression from H&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods operate at the spot level, where signals from multiple cells are aggregated and critical cellular heterogeneity is obscured. Extending this paradigm to single-cell resolution is non-trivial. Naively applying pathology foundation models faces a scale mismatch: their patch-level representations mix multiple cells, whereas per-cell cropping or resizing distorts morphology and removes local context. Conversely, segmentation-based models without strong pretrained visual encoders often lack the morphological representation capacity needed for accurate molecular prediction and inherit errors from imperfect cell boundary masks. Here, we present CELLO, an efficient end-to-end framework that performs a single pathology foundation model forward pass per image and uses grid sampling to extract location-specific features for all cells simultaneously. We further introduce a distance-decay cross-attention module that refines each cell representation using spatially biased local morphological context. Using 52 public Xenium-H&E pairs from HEST-1k that span 12 organs and approximately 10 million cells, CELLO improves the average predictive accuracy over the evaluated baselines while reducing the mean whole-slide inference time compared to DeepSpot2Cell, a 14.0x speed-up on average that excludes upstream cell segmentation. Our work establishes a scalable foundation for single-cell gene expression prediction from H&E images.
发表机构
- The Chinese University of Hong Kong(香港中文大学)
- Stanford University School of Medicine(斯坦福大学医学院)
机构由 AI 辅助整理,请以论文原文为准。