arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VOICE:一种融合原位单细胞基因表达直接预测与基于检索预测的视觉-组学基础模型

VOICE: A Vision-Omics Foundation Model Integrating Direct and Retrieval-Based Prediction of In-situ Single-Cell Gene Expression

Xin Luo, Yicheng Tao, Haoxuan Zeng, Suyuan Wang, Chenzi Ouyang, Meiqi Zhu, Kai Liu, Shuibing Chen, Jie Liu

arXiv 2608.08366首次发表:更新:

发表机构

University of Michigan; Weill Cornell Medicine(密歇根大学; 威尔康奈尔医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出多模态基础模型 VOICE,通过对比学习对齐 H&E 形态与单细胞表达嵌入,融合直接回归与参考检索分支,泛化性良好且在七项指标上优于现有方法,可从 H&E 图像预测单细胞基因表达。

AI 中文摘要

空间转录组学可解析单细胞分辨率的基因表达,但成本高昂,仅能覆盖数百至数千个基因的靶向面板,且仅适用于少量样本。相比之下,H&E 成像成本低廉且被大规模常规采集,这使得从形态学直接预测单细胞表达成为将分子分析应用于大型组织档案的可行方法。因此,我们提出 VOICE,这是一种利用配对 Xenium 数据从 H&E 图像预测单细胞基因表达的多模态基础模型。VOICE 首先通过对比学习在 2300 万个细胞上进行训练,将病理学基础模型输出的以细胞为中心的 H&E 形态与转录组基础模型输出的单细胞表达嵌入进行对齐。接下来,它通过两个分支预测表达:一个分支从形态学直接回归表达,另一个分支从相似的参考细胞中检索已测量的表达,以恢复不具备形态学信号的基因。由于不同基因的形态可预测性存在差异,VOICE 采用按基因分配权重的方式融合两个分支。训练完成后,VOICE 可泛化到保留的患者、切片以及 Xenium 中部分重叠的基因面板,且在七项指标上始终优于现有的单细胞表达预测方法。

英文摘要

Spatial transcriptomics can resolve gene expression at single-cell resolution, but it is costly, limited to targeted panels of a few hundred to a few thousand genes, and applicable to only a small number of samples. H&E imaging, by contrast, is cheap and collected routinely at scale. This makes predicting single-cell expression directly from morphology a practical way to bring molecular analysis to large tissue archives. We therefore present VOICE, a multimodal foundation model that predicts single-cell gene expression from H&E images using paired Xenium data. VOICE first aligns cell centered H&E morphology from a pathology foundation model with single-cell expression embeddings from a transcriptome foundation model, trained using contrastive learning over 23 million cells. Next it predicts expression through two branches. One branch directly regresses expression from morphology. The other branch retrieves measured expression from similar reference cells, recovering genes that do not have morphological signal. Because genes vary in morphological predictability, VOICE fuses the two branches with a per-gene weight. After training, VOICE generalizes to heldout patients, slides, and partially overlapping gene panels from Xenium, and it consistently outperforms prior single-cell expression prediction methods on seven metrics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑