arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解耦共享表示可改进形态-转录组整合

Disentangled Shared Representations Improve Morpho-Transcriptomic Integration

Julian Ostermaier, Swann Ruyter, Reuben Dorent, Daniel Racoceanu

arXiv 2608.14355首次发表:更新:

发表机构

ESPCI Paris, PSL University; Sorbonne Université; CNRS; Inserm; AP-HP; Inria; Paris Brain Institute – ICM(巴黎高等物理化工学院,PSL大学; 索邦大学; 法国国家科学研究中心; 法国国家健康与医学研究院; 巴黎公共医疗集团; 法国国家信息与自动化研究所; 巴黎脑研究所 – ICM)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对空间转录组学的多模态表示学习问题,对比不同模型的标准与解耦变体,发现解耦共享与模态特有信息可提升相关指标,为多模态学习提供评估框架。

AI 中文摘要

空间转录组学(Spatial transcriptomics, ST)可同时分析基因表达与组织形态,为学习捕获形态-转录组共享结构的多模态表示提供了可能。但标准多模态模型常将各模态压缩至共同隐空间,未显式分离共享与模态特有的变异源,可能限制下游效用。本研究探究显式解耦共享与私有隐成分是否能改进配对苏木精-伊红(Hematoxylin & Eosin, H&E)染色图像与ST数据的多模态表示学习。在匹配的实验条件下,于两个癌症队列中对比基于变分自编码器(VAE)的方法与对比方法的标准变体和解耦变体,通过跨模态重构、下游探测及跨模态探测迁移对表示进行评估。实验揭示两个主要趋势:其一,对比目标函数的下游探测性能优于基于VAE的模型;其二,解耦变体可提升选定的重构与探测指标,尽管提升幅度取决于模型家族、任务、方向和解耦强度。总体而言,本研究结果表明,显式分解共享与模态特有信息可改进空间转录组学的多模态表示学习,并为未来基础模型提供有用的评估框架。

英文摘要

Spatial transcriptomics (ST) enables the simultaneous profiling of gene expression and tissue morphology, creating an opportunity to learn multimodal representations capturing shared morpho-transcriptomic structure. However, standard multimodal models often compress modalities into a common latent space without explicitly separating shared and modality-specific sources of variation, which may limit downstream utility. We investigate whether explicit disentanglement of shared and private latent components improves multimodal representation learning for paired Hematoxylin \& Eosin (H\&E) and ST data. We compare VAE-based and contrastive approaches, each in standard and disentangled variants, across two cancer cohorts under matched experimental conditions. Representations are evaluated using cross-modal reconstruction, downstream probing and cross-modal probe transfer. The experiments suggest two main trends. First, contrastive objectives yield higher downstream probing performance than VAE-based models. Second, disentangled variants improve the selected reconstruction and probing metrics, although the gains depend on the model family, task, direction, and disentanglement strength. Overall, our results suggest that explicitly factorizing shared and modality-specific information can improve multimodal representation learning for spatial transcriptomics and provides a useful evaluation framework for future foundation models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑