基于28万份常规报告的肠镜报告驱动的视觉-语言基础模型
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
浏览论文内容
中文总结 AI 辅助
该研究开发了肠镜视觉-语言基础模型EndoCLIP,利用28万份常规肠镜记录恢复的图像-文本对训练,在多项任务中优于通用模型,良恶性分类性能接近专家,可实现临床目标的语言指定。
中文摘要 AI 辅助
尽管常规肠镜报告中记录了丰富的专家描述,但视觉-语言模型在肠镜检查中的应用仍不充分。这些报告记录了病变的外观、大小和位置,但总结的是整个检查过程而非单个帧的描述,导致临床发现与对应图像的关联较弱。本文开发了EndoCLIP,这是一种肠镜视觉-语言基础模型,通过从280476份常规肠镜记录中逐步恢复出125756个病变级图像-文本对进行训练。在病变级图像-文本检索、结构化报告生成以及6项多中心临床分类任务中,EndoCLIP在零样本和线性探测设置下均优于通用和生物医学视觉-语言编码器。在良恶性分类任务中,其线性探测性能在涉及12名内镜医师的盲法研究中接近专家阅片者的水平。这些结果表明,恢复发现与帧的对应关系可将常规文档转化为可扩展的监督信号,使临床目标能以语言指定,无需为每个任务单独标注。
英文摘要
Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records. Across lesion-level image-text retrieval, structured report generation and six multi-centre clinical classification tasks, EndoCLIP outperforms general-purpose and biomedical vision-language encoders in both zero-shot and linear-probe settings. On benign-versus-malignant classification, its linear probe approaches the performance of expert readers in a blinded study involving 12 endoscopists. These results suggest that recovering finding-to-frame correspondence can transform routine documentation into scalable supervision, enabling clinical targets to be specified in language rather than separately annotated for each task.
发表机构
- Digital Medical Research Center, School of Basic Medical Sciences, Fudan University(复旦大学基础医学院数字医学研究中心)
- Shanghai Collaborative Innovation Center of Endoscopy(上海内镜诊疗协同创新中心)
- Zhejiang University(浙江大学)
- Shanghai Institute for Advanced Study of Zhejiang University(浙江大学上海高等研究院)
- Alliance Manchester Business School, The University of Manchester(曼彻斯特大学联盟曼彻斯特商学院)
- Data Science Institute, Imperial College London(伦敦帝国理工学院数据科学研究所)
- Microsoft Research Asia(微软亚洲研究院)
机构由 AI 辅助整理,请以论文原文为准。