基于概念的H&E图像基因表达预测解释
Concept-based explanation of gene expression prediction from H&E images
浏览论文内容
中文总结 AI 辅助
本研究提出结合相关性传播与概念发现的可解释ViT框架,用于从H&E图像预测空间转录组学,可关联分子表型与组织形态,能准确预测结直肠癌相关ST特征并分层患者结局。
中文摘要 AI 辅助
病理学基础模型的最新进展已能从常规H&E图像中准确预测空间转录组学(ST)。然而,现有针对视觉Transformer(ViT)模型的可解释性方法大多局限于局部热图,无法揭示形态学概念如何影响ST预测。本文提出一种结合相关性传播与概念发现的可解释框架,将转录程序与组织形态学关联起来。我们开发了一种基于ViT的H&E图像虚拟ST预测框架,该框架结合了感知ViT的逐层相关性传播与松弛原型TopK稀疏自编码器的概念发现,既能提供局部解释,又能为与转录程序相关的形态模式提供全局见解。我们将该框架应用于HEST-1k队列的结直肠癌ST数据,并在TCGA COAD中评估其泛化性。我们的架构能准确预测临床相关的ST特征及伴随的分子表型,实测与预测基因表达谱揭示了大量样本中结直肠癌亚型iCMS2与iCMS3的显著空间异质性;空间分辨率及聚合iCMS分类的加权F1分数分别为0.872、0.819(TCGA COAD中为0.770),且二者均可对患者结局进行分层。除预测外,我们的框架还构建了基于相关性的概念图谱,将分子表型与组织病理学表征关联;对比激活与相关性衍生的概念可知,相关性能更直接地连接组织形态与下游预测。我们确立了空间预测的概念型解释通用策略,该框架可便捷应用于多种基于ViT的病理学模型。
英文摘要
Recent advances in pathology foundation models have enabled accurate prediction of spatial transcriptomics (ST) from routine H&E images. However, existing explainability methods for vision transformer (ViT)-based models are largely limited to local heatmaps and do not reveal how morphological concepts contribute to ST predictions. Here, we introduce an explainable framework that combines relevance propagation and concept discovery to link transcriptional programs to tissue morphology. We developed a ViT-based framework for virtual ST from H&E images that combines ViT-aware layer-wise relevance propagation with relaxed archetypal TopK sparse autoencoder-based concept discovery. This approach provides both local explanations and global insights into the morphological patterns associated with transcriptional programs. We applied the framework to colorectal cancer ST data from the HEST-1k cohort and evaluated its generalizability in TCGA COAD. Our architecture accurately predicts clinically relevant ST signatures and accompanying molecular phenotypes. Measured and predicted gene expression profiles reveal substantial spatial heterogeneity of the colorectal cancer subtypes iCMS2 and iCMS3 across a large number of samples. Spatially resolved and aggregated iCMS classification achieve weighted F1 scores of 0.872 and 0.819 (0.770 in TCGA COAD), respectively, and both stratify patient outcome. Beyond prediction, our framework establishes a relevance-based concept atlas linking molecular phenotypes to histopathological representations. Comparison of activation- with relevance-derived concepts demonstrates that relevances provide a more direct link between tissue morphology and downstream predictions. We establish a general strategy for concept-based explanation of spatial prediction, and our framework is readily applicable to a broad range of ViT-based pathology models.
发表机构
- Institute of Pathology – Charité Universitätsmedizin(夏里特医学院病理学研究所)
- Fraunhofer Heinrich-Hertz-Institute(弗劳恩霍夫海因里希·赫兹研究所)
- Center for Computational Biology (CBIO), Mines Paris, PSL University(巴黎矿业学院计算生物学中心(巴黎文理研究大学))
- Institut Curie, PSL University(居里研究所(巴黎文理研究大学))
- German Cancer Research Center (DKFZ)(德国癌症研究中心)
- Technological University Dublin(都柏林理工大学)
- Technische Universität Berlin(柏林工业大学)
- BIFOLD – Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究所)
机构由 AI 辅助整理,请以论文原文为准。