LoMeVQA:纵向医学视觉问答综合基准
LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA
浏览论文内容
中文总结 AI 辅助
该研究提出纵向医学视觉问答基准LoMeVQA,发现现有多模态大语言模型在该任务上时间推理能力不足,推出MedLong-8B实现最优性能,并开展相关分析。
中文摘要 AI 辅助
在临床实践中,患者常需在多次随访中接受多项影像学检查,从而产生纵向数据。对这类时间信息进行建模,对于可靠评估疾病进展和治疗应答至关重要。然而,尽管多模态大语言模型(MLLMs)发展迅速,纵向医学视觉推理仍在很大程度上未被探索。为填补这一空白,我们提出LoMeVQA,这是一个包含20.6万对纵向视觉问答(VQA)的综合基准,用于时间医学图像分析。LoMeVQA涵盖五项任务:进展分类、进展描述、进展报告生成、差异区域定位及差异区域描述。为构建该数据集,我们开发了一条自动化流程,该流程:(1)按时间顺序整理患者记录;(2)通过医学知识图谱提取具有临床意义的实体;(3)对实体的时间演变进行建模,以指导大语言模型生成高质量的纵向VQA对。大量评估表明,通用型和医学领域的MLLMs在LoMeVQA上的表现均不佳,暴露了其在时间推理方面的显著局限。为解决这些局限,我们推出MedLong-8B,该模型在所有任务中均达到了最先进的性能。除基准测试外,我们还开展了详细分析,揭示了关键失败模式,并阐明了如何改进纵向医学视觉推理。我们的数据可通过以下网址获取:this https URL
英文摘要
In clinical practice, patients often undergo multiple imaging examinations over successive visits, yielding longitudinal data. Modeling such temporal information is crucial for reliable assessment of disease progression and treatment response. However, despite the rapid advancement of multimodal large language models (MLLMs), longitudinal medical visual reasoning remains largely underexplored. To fill this gap, we propose LoMeVQA, a comprehensive benchmark consisting of 206K longitudinal visual question answering (VQA) pairs for temporal medical image analysis. LoMeVQA covers five tasks: progress classification, progress description, progress report generation, differential region grounding, and differential region description. To construct the dataset, we develop an automated pipeline that (1) organizes patient records chronologically, (2) extracts clinically meaningful entities via a medical knowledge graph, and (3) models their temporal evolution to guide large language models in generating high-quality longitudinal VQA pairs. Extensive evaluations demonstrate that both general-purpose and medical-domain MLLMs perform poorly on LoMeVQA, revealing substantial limitations in temporal reasoning. To address these limitations, we introduce MedLong-8B, which achieves state-of-the-art performance across all tasks. Beyond benchmarking, we conduct detailed analyses that uncover key failure modes and shed light on how to improve longitudinal medical visual reasoning. Our data is available at: https://github.com/pepperbubble/LoMeVQA
发表机构
- Tongji University(同济大学)
- East China Normal University(华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。