利用大规模语言和视觉模型从大规模图像-文本结肠镜检查记录中进行知识提取与蒸馏
Knowledge Extraction and Distillation from Large-Scale Image-Text Colonoscopy Records Leveraging Large Language and Vision Models
- School of Basic Medical Sciences, Fudan University(复旦大学基础医学院)
- Shanghai Key Laboratory of MICCAI(上海市MICCAI重点实验室)
- Data Science Institute, Imperial College London(伦敦帝国理工学院数据科学研究所)
- Shanghai Collaborative Innovation Centre of Endoscopy(上海市内镜微创技术协同创新中心)
- Endoscopy Centre and Endoscopy Research Institute, Zhongshan Hospital, Fudan University(复旦大学附属中山医院内镜中心及内镜研究所)
- Academy for Engineering and Technology, Fudan University(复旦大学工程与技术学院)
- School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
- Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程学系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对结肠镜分析中数据标注昂贵的问题,提出EndoKED范式,利用大语言和视觉模型自动从大规模图像-文本记录中提取知识并生成像素级标注,验证表明其训练出的模型性能优越且可泛化。
AI中文摘要:
人工智能系统在结肠镜检查分析中的开发通常需要专家标注的图像数据集。然而,数据集规模和多样性的限制阻碍了模型性能和泛化能力。来自常规临床实践的图像-文本结肠镜检查记录,包含数百万张图像和文本报告,是宝贵的数据来源,尽管对其进行标注非常耗费人力。在此,我们利用大规模语言和视觉模型的最新进展,提出了EndoKED,一种用于深度知识提取和蒸馏的数据挖掘范式。EndoKED自动化地将原始结肠镜检查记录转化为具有像素级标注的图像数据集。我们使用多中心原始结肠镜检查记录数据集(约100万张图像)验证了EndoKED,证明了其在训练息肉检测和分割模型方面的优越性能。此外,EndoKED预训练的视觉骨干网络能够实现光学活检的数据高效且可泛化的学习,在回顾性和前瞻性验证中均达到了专家级性能。
英文摘要:
The development of artificial intelligence systems for colonoscopy analysis often necessitates expert-annotated image datasets. However, limitations in dataset size and diversity impede model performance and generalisation. Image-text colonoscopy records from routine clinical practice, comprising millions of images and text reports, serve as a valuable data source, though annotating them is labour-intensive. Here we leverage recent advancements in large language and vision models and propose EndoKED, a data mining paradigm for deep knowledge extraction and distillation. EndoKED automates the transformation of raw colonoscopy records into image datasets with pixel-level annotation. We validate EndoKED using multi-centre datasets of raw colonoscopy records (~1 million images), demonstrating its superior performance in training polyp detection and segmentation models. Furthermore, the EndoKED pre-trained vision backbone enables data-efficient and generalisable learning for optical biopsy, achieving expert-level performance in both retrospective and prospective validation.