EndoVLM:一种通过解剖学引导的稀疏性与渐进式对齐实现的内窥镜视觉-语言预训练模型
EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment
浏览论文内容
中文总结 AI 辅助
本研究提出了EndoVLM这一内窥镜视觉-语言预训练模型,通过解剖学引导的稀疏池化、渐进式语义感知对齐等技术,在34.8万例内窥镜检查数据上预训练,其性能优于现有基础模型且具备零样本泛化能力。
中文摘要 AI 辅助
基础模型(FMs)的发展对推进内窥镜图像分析至关重要。然而,现有的内窥镜基础模型主要依赖于对单模态图像或视频的自监督学习,忽略了临床报告中包含的丰富语义知识。此外,由于存在根本性的模态差距,有效利用这些记录受到阻碍:结构化的解剖学描述无法自然地映射到高冗余、未整理的视觉流中的特定帧。在本文中,我们提出了EndoVLM,这是一种新型的视觉-语言基础模型,基于超过34.8万例内窥镜检查进行预训练,每例检查都包含一份临床报告及其对应的图像集合。解剖学引导的稀疏池化机制利用文本描述作为查询来驱动稀疏注意力,将语义显著的帧高效聚合为跨冗余图像集的解剖学特定视觉表示。接下来,渐进式语义感知对齐策略通过结构化软目标对临床分类(解剖学和病理状态)进行建模,弥合了从全局患者级匹配到细粒度局部对齐的差距。最后,语义集中式掩码自动编码器仅应用于这些语义丰富的帧,将低级视觉精度与鲁棒的高级语义表示相结合。在各种下游任务上进行的大量实验表明,EndoVLM的性能优于现有的基础模型,并且与特定任务方法相比具有竞争力。值得注意的是,EndoVLM还表现出鲁棒的零样本泛化能力,凸显了其在更广泛临床应用中的潜力。
英文摘要
The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely on self-supervised learning from uni-modal images or videos, overlooking the rich semantic knowledge contained in clinical reports. Furthermore, effectively leveraging these records is hindered by a fundamental modality gap: structured anatomical descriptions are not naturally mapped to specific frames within the high-redundancy, uncurated visual streams. In this paper, we present EndoVLM, a novel vision-language FM pre-trained on over 348K endoscopic examinations, each pairing a clinical report with its corresponding image collection. An Anatomy-Guided Sparse Pooling mechanism utilizes textual descriptions as queries to drive sparse attention, efficiently aggregating semantically salient frames into anatomy-specific visual representations across redundant image-sets. Next, a Progressive Semantic-Aware Alignment strategy models clinical taxonomy (anatomy and pathological status) via structured soft targets, bridging the gap from global patient-level matching to fine-grained localized alignment. Finally, a Semantic-Concentrated Masked Autoencoder is applied exclusively to these semantic-rich frames, integrating low-level visual precision with robust high-level semantic representation. Extensive experiments across various downstream tasks demonstrate that EndoVLM outperforms existing foundation models and remains competitive with task-specific methods. Remarkably, EndoVLM also exhibits robust zero-shot generalization capabilities, highlighting its potential for broader clinical application.
发表机构
- DAMO Academy, Alibaba Group(阿里巴巴达摩院)
- The First Affiliated Hospital of Zhejiang Chinese Medical University(浙江中医药大学附属第一医院)
- Shanghai Jiao Tong University(上海交通大学)
- Hupan Lab(湖畔实验室)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。