arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CT-ΔBench:基于视觉语言模型的纵向3D医学影像差异报告基准

CT-$Δ$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

Kegeng Tang, Jingbo Wang, Shaogang Ren, Zihao Wang

arXiv 2608.11534首次发表:更新:

发表机构

University of Tennessee at Chattanooga(田纳西大学查塔努加分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对现有医学基础模型缺乏纵向影像差异处理能力的问题,构建了专用基准CT-ΔBench,开发感知变化指标与医师验证流程,对比不同推理方式并提出基线模型DeltaMed,为时间感知医学基础模型奠定基础。

AI 中文摘要

在医学影像领域,计算机断层扫描(CT)的临床价值不仅在于描绘当前疾病状态,更关键的是能对连续扫描进行纵向比较以判断疾病进展,这一过程是疗效评估、复发检测及患者持续管理的基础。尽管时间维度的对比在临床决策中具有核心作用,但现有医学基础模型大多局限于单一时段研究的理解,导致基于时间维度的交叉检查问题未得到充分解决。为填补这一空白,本文研究纵向影像差异报告任务,该任务要求模型获取同一患者的两次时间间隔扫描,生成描述两者间间隔变化的临床意义报告。本文提出针对该任务的专用基准CT-ΔBench,采用患者级划分以防止信息泄露;为超越表面文本相似性更好评估该任务,进一步开发专门捕捉临床有意义纵向变化的感知变化指标,并开展独立医师验证以评估合成参考及事件提取流程的可靠性;还对比了直接配对CT推理与先生成单时间点报告再进行文本差异分析的间接两阶段流程;最后提出用于直接配对CT差异报告的基线模型DeltaMed,并在基准训练集上进行训练。上述工作为能更好反映真实世界纵向临床推理的时间感知医学基础模型奠定了基础。

英文摘要

In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitudinal comparison of serial scans to determine disease evolution, a process that underpins response assessment, recurrence detection, and ongoing patient management. Yet, despite this central role of temporal comparison in clinical decision-making, existing medical foundation models remain largely confined to single-study understanding, leaving temporally grounded cross-examination insufficiently addressed. To address this gap, we study longitudinal imaging difference reporting, a task in which a model takes two temporally separated scans from the same patient and generates a clinically meaningful report describing interval changes between them. We introduce CT-$Δ$Bench, a dedicated benchmark for this task with patient-level splitting to prevent information leakage. To better evaluate this task beyond surface-level text similarity, we further develop change-aware metrics specifically designed to capture clinically meaningful longitudinal changes, and conduct an independent physician validation to assess the reliability of the synthesized references and event extraction pipeline. We also compare direct paired-CT reasoning with an indirect two-stage pipeline that first generates single-timepoint reports and then performs textual differencing. Finally, we propose DeltaMed, a baseline model for direct paired-CT difference reporting, and train it on the benchmark training set. Together, these contributions lay the groundwork for temporally aware medical foundation models that better reflect real-world longitudinal clinical reasoning.

CommentsAccepted by COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑