arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

计算病理学中用于全模态病理切片表征学习的协同信息解缠

Synergistic Information Disentanglement for Omni-modal Slide Representation Learning in Computational Pathology

Mingxin Liu, Chengfei Cai, Anwen Lu, Pengbo Xu, Jun Li, Jinze Li, Depin Chen, Jun Xu

arXiv 2609.02118首次发表:更新:

发表机构

Nanjing University of Information Science and Technology; Taizhou University; Harbin Medical University(南京信息工程大学; 泰州大学; 哈尔滨医科大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对计算病理学全模态切片表征学习的模态坍塌问题,提出Φ-Omni协同信息解缠框架,经多数据集预训练后在少样本任务中优于基线方法。

AI 中文摘要

在计算病理学(CPath)领域,开发整合组织学、基因组学与临床报告的全模态自监督学习(SSL)模型,可为全切片图像(WSI)提供可迁移的表征学习能力。现有方法通过对比对齐将异质模态隐含地强制映射到统一隐空间,导致模态坍塌,即独特的协同诊断信号(记为Φ)被舍弃,转而保留琐碎冗余信息。我们假设,最强的任务无关SSL训练信号源于对协同交互的提炼,而非仅对齐共享冗余。为此,我们提出基于部分信息分解(PID)理论的协同信息解缠框架——Φ-Omni,用于切片表征学习。与标准对比方法不同,Φ-Omni采用由所提ΦID目标调控的协同信息瓶颈(SIB),该目标明确抑制边缘冗余,同时最大化不可约协同性,从而提炼高阶跨模态交互。在乳腺癌队列(n=1031)与肺癌队列(n=919)上进行预训练后,Φ-Omni在涵盖8项任务的5个独立外部数据集上,相较于监督与SSL基线方法,展现出更优的少样本性能。源代码可在此获取。

英文摘要

In computational pathology (CPath), developing omni-modal self-supervised learning (SSL) models that integrate histology, genomics, and clinical reports enables transferable representation learning for whole slide images (WSIs). Existing approaches implicitly force heterogeneous modalities into a uniform latent space by contrastive alignment, causing modality collapse where unique, synergistic diagnostic signals (termed as $\mathrmΦ$) are discarded in favor of trivial redundancy. We hypothesize that the strongest task-agnostic SSL training signal stems from distilling the synergistic interactions over merely aligning shared redundancy. To this end, we introduce \textsc{$\mathrmΦ$-Omni}, a synergistic information disentanglement framework grounded in Partial Information Decomposition (PID) theory for slide representation learning. Unlike standard contrastive approaches, \textsc{$\mathrmΦ$-Omni} employs a Synergistic Information Bottleneck (SIB) regulated by the proposed $\mathrmΦ\text{ID}$ objective, which explicitly suppresses marginal redundancy while maximizing irreducible synergy, thereby distilling high-order cross-modal interactions. Following pretraining on breast ($n$=1031) and lung ($n$=919) cohorts, \textsc{$\mathrmΦ$-Omni} demonstrates superior few-shot performance across five independent external datasets spanning eight tasks compared to supervised and SSL baselines. Source code is available here.

CommentsNeeds further revision

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑