arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

实现多模态卒中复发预测的视觉与跨模态学习:一种可解释的两步框架

Enabling Vision and Cross-Modal Learning for Multimodal Stroke Recurrence Prediction: An Interpretable Two-Step Framework

Christian Gapp, Elias Tappeiner, Martin Welk, Karl Fritscher, Stephanie Mangesius, Constantin Eisenschink, Philipp Deisl, Michael Knoflach, Astrid E. Grams, Elke R. Gizewski, Rainer Schubert

arXiv 2609.22271首次发表:更新:

发表机构

UMIT TIROL – Private University for Health Sciences and Health Technology; VASCage – Centre on Clinical Stroke Research; Medical University of Innsbruck(UMIT TIROL – 健康科学与健康技术私立大学; VASCage – 临床卒中研究中心; 因斯布鲁克医科大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种可解释的两步框架,通过自监督预训练和策略性微调,实现多模态卒中复发预测中视觉与跨模态学习的有效整合,克服模态不平衡并提升预测性能。

AI 中文摘要

多模态卒中复发预测需要有效整合异质的临床和影像数据,然而模态不平衡常常导致模型过度依赖主导模态,而未能充分利用互补信息。尽管自监督预训练和选择性参数冻结通常被用于改进表示学习和微调稳定性,但它们对多模态医学模型中模态贡献和跨模态行为的影响在很大程度上仍未得到探索。在本工作中,我们研究了在3D CTA扫描上的图像预训练是否能减少模态不平衡并改善卒中复发预测(我们最近处理的一项临床关键任务)的跨模态整合。为此,两个多模态神经网络以自监督方式预训练,随后使用两种不同的冻结策略进行微调。它们的性能和模态利用情况与我们先前工作中的基线模型以及本研究从头训练的模型进行了比较。我们的结果表明,自监督预训练能够更有效地利用多模态图像-表格数据集,优于先前的基线和所有非预训练模型。值得注意的是,表现最佳的基于视觉Transformer的神经网络成功克服了单模态崩溃。协同分析揭示了视觉与性别和冠心病之间的显著交互作用,提示了与卒中复发相关的临床模式。总体而言,我们的发现表明,自监督预训练和策略性微调支持更平衡的模态利用,并实现有意义的跨模态交互。代码可在以下https URL公开获取。

英文摘要

Multimodal stroke recurrence prediction requires effective integration of heterogeneous clinical and imaging data, yet modality imbalance often causes models to over-rely on dominant modalities and underutilize complementary information. While self-supervised pretraining and selective parameter freezing are commonly employed to improve representation learning and fine-tuning stability, their effect on modality contributions and cross-modal behavior in multimodal medical models remains largely unexplored. In this work, we investigate whether image pretraining on 3D CTA scans reduces modality imbalance and improves cross-modal integration for stroke recurrence prediction, a clinically critical task we recently addressed. To this end, two multimodal neural networks are pretrained in a self-supervised manner and subsequently fine-tuned using two distinct freezing strategies. Their performance and modality utilization are compared against both the baseline model from our previous work and models trained entirely from scratch in this study. Our results demonstrate that self-supervised pretraining enables more effective utilization of the multimodal image-tabular dataset, outperforming both the prior baseline and all non-pretrained models. Notably, the best-performing Vision Transformer based neural network successfully overcomes unimodal collapse. Synergy analysis reveals significant interactions between vision and both gender and CHD, suggesting clinically relevant patterns for stroke recurrence. Overall, our findings demonstrate that self-supervised pretraining and strategic fine-tuning support more balanced modality utilization and enable meaningful cross-modal interactions. Code is publicly available at https://github.com/ChristianGappGit/SSL_Pretraining.

CommentsML-CDS 2026: Multimodal Learning and Fusion Across Scales for Clinical Decision Support, MICCAI 2026, Strasbourg, France

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑