MedTVL:利用视觉与语言进行医学时间序列分类
MedTVL: Harnessing Vision and Language for Medical Time Series Classification
浏览论文内容
中文总结 AI 辅助
该研究提出文本引导双路径架构MedTVL,结合卷积时间路径与Transformer视觉路径,辅以自适应医学文本语义和混合专家机制,支持多模态对比学习,在多医学任务上展现出优越的分类性能与可迁移性。
中文摘要 AI 辅助
医学时间序列(MedTS)分类的多模态学习近期进展凸显了整合互补模态对临床决策的益处,但现有方法通常聚焦于双模态交互(如时间序列与文本),尚未充分探索时间序列、视觉与语言三者间的协同作用。受诊断实践中数值评估、视觉检查与临床背景协同的启发,我们提出MedTVL——一种专为MedTS分类设计的文本引导双路径架构。具体而言,该架构协同了基于卷积的时间路径,用于从原始数值序列中提取细粒度时间动态;以及基于Transformer的视觉路径,用于从时间序列衍生图像中提取整体形态结构。这种跨模态与架构异质性的结合提供了全面的诊断视角。为进一步解决潜在的诊断歧义,两条路径均受自适应医学文本语义引导。最后,混合专家(Mixture-of-Experts)机制将每个样本动态路由至专门的融合专家,捕捉样本对时间路径与视觉路径输出的特定依赖。此外,MedTVL支持多模态对比学习,以缓解临床标签稀缺的挑战。在多个医学数据集和任务(涵盖监督、少样本及对比学习设置)上开展的大量实验,证明了MedTVL的优越性与可迁移性,凸显了其在稳健临床决策支持方面的潜力。
英文摘要
Recent advancements in multimodal learning for medical time series (MedTS) classification highlight the benefits of integrating complementary modalities for clinical decision. However, existing methods typically focus on bi-modal interactions (e.g., time series and text), leaving the tri-modal synergy between time series, vision, and language largely unexplored. Inspired by diagnostic practice synergizing numerical assessment, visual inspection and clinical context, we introduce MedTVL, a text-guided dual-pathway architecture tailored for MedTS classification. Specifically, it synergizes a convolution-based temporal pathway for fine-grained temporal dynamics from raw numerical sequences and a transformer-based visual pathway for holistic morphological structures from time-series-derived images. Such combination of cross-modal and architectural heterogeneity provides a comprehensive diagnostic perspective. To further resolve potential diagnostic ambiguity, both pathways are guided by adaptive medical textual semantics. Finally, a Mixture-of-Experts mechanism dynamically routes each instance to specialized fusion experts, capturing instance-specific reliance on the temporal and visual pathway outputs. In addition, MedTVL supports multimodal contrastive learning to mitigate the clinical label scarcity challenge. Extensive experiments across multiple medical datasets and tasks, spanning supervised, few-shot, and contrastive learning settings, demonstrate the superiority and transferability of MedTVL, highlighting its potential for robust clinical decision support.
发表机构
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。