arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34439cs.CV

临床轨迹对齐用于医学视觉-语言预训练

Clinical Trajectory Alignment for Medical Vision-Language Pre-training

  • Institute of Intelligent Information Processing, Shanxi University(山西大学智能信息处理研究所)
  • Alliance Manchester Business School, The University of Manchester(曼彻斯特大学联盟曼彻斯特商学院)

机构由 AI 辅助整理,请以论文原文为准。

Huimin Yan, Xian Yang, Zhi Wang, Liang Bai

AI总结:

MedCTA将医学视觉-语言预训练从就诊级匹配转向学习临床变化,通过异常和患者病程双范围建模,结合LLM提取趋势监督,在时间分类、检索和零样本任务上超越基线。

AI中文摘要:

医学视觉-语言预训练主要遵循就诊级别的图像-报告匹配范式,在个体就诊时对齐配对的图像和报告。虽然这种方法对静态跨模态对应有效,但该范式对纵向临床变化(例如异常是否随时间改善、保持稳定或恶化)提供的监督有限。学习这种变化具有挑战性,因为时间语义隐含在自由文本报告中,且同一患者体内的不同异常可能异步演变,甚至朝相反方向发展。我们提出MedCTA,将医学视觉-语言预训练从就诊级别的跨模态匹配重构为学习临床变化。MedCTA不是将患者病史压缩为单一时间表示,而是在两个互补范围内建模临床变化。在异常范围内,临床引导的查询构建异常条件下的视觉和文本轨迹,以捕捉异质性异常演变。在患者病程范围内,对全局图像和报告序列进行建模,以捕捉超越任何个体异常的整体临床进展。结构化趋势监督由离线LLM解析器从纵向报告中提取,无需手动时间标注。结合静态图像-报告对齐,MedCTA学习到的表示既保留就诊级别的跨模态对应,又编码纵向变化语义。在时间图像分类、图像-文本检索和零样本分类上的实验显示,相对于强医学视觉-语言基线,取得了一致的改进。

英文摘要:

Medical vision-language pre-training largely follows a visit-level image-report matching paradigm, aligning paired images and reports at individual visits. While effective for static cross-modal correspondence, this paradigm provides limited supervision for longitudinal clinical change, such as whether abnormalities improve, remain stable, or worsen over time. Learning such change is challenging because temporal semantics are implicit in free-text reports, and different abnormalities within the same patient may evolve asynchronously or even in opposite directions. We propose MedCTA, which reframes medical vision-language pre-training from visit-level cross-modal matching to learning clinical change. Rather than compressing a patient history into a single temporal representation, MedCTA models clinical change at two complementary scopes. At the abnormality scope, clinically grounded queries construct abnormality-conditioned visual and textual trajectories to capture heterogeneous abnormality evolution. At the patient-course scope, global image and report sequences are modeled to capture overall clinical progression beyond any individual abnormality. Structured trend supervision is extracted from longitudinal reports by an offline LLM parser, removing the need for manual temporal annotations. Combined with static image-report alignment, MedCTA learns representations that preserve visit-level cross-modal correspondence while encoding longitudinal change semantics. Experiments on temporal image classification, image-text retrieval, and zero-shot classification show consistent gains over strong medical vision-language baselines.

↑