发表机构
Zhejiang University; Shanghai Institute for Advanced Study, Zhejiang University; Shanghai Key Laboratory of MICCAI; Digital Medical Research Center, School of Basic Medical Sciences, Fudan University; Endoscopy Center and Endoscopy Research Institute, Zhongshan Hospital, Fudan University; Shanghai Collaborative Innovation Center of Endoscopy; Alliance Manchester Business School, The University of Manchester(浙江大学; 浙江大学上海高等研究院; 上海市MICCAI重点实验室; 复旦大学基础医学院数字医学研究中心; 复旦大学附属中山医院内镜中心及内镜研究所; 上海市内镜协同创新中心; 曼彻斯特大学联盟曼彻斯特商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对内窥镜息肉报告需求,提出一种不修改预训练权重的上下文融合框架,通过隐式指令与显式转换上下文对冻结通用VLM专用化,在2056张标注图像实验中实现最优性能且参数量仅为冻结VLM的0.006%。
AI 中文摘要
可靠的内窥镜息肉报告需要在单一记录中整合病灶尺寸量化、标准化巴黎分类以及具有临床意义的形态学描述。通用视觉语言模型(VLM)为图像理解与报告生成提供了统一接口。然而,现有的专用策略通常依赖于特定任务模型或模型权重适配,如何在保留该统一接口和VLM预训练能力的同时引入可靠的专业知识,这一问题尚未得到解决。我们提出一种上下文融合框架,该框架通过隐式指令上下文和显式转换上下文对冻结的通用VLM进行专用化,且不修改其预训练权重。具体而言,自监督息肉编码器检索相关的图像-报告对作为显式的、查询特定的证据,而学习到的连续专业令牌则提供跨案例共享的隐式指令上下文。实验在2056张专家标注的公开内窥镜图像上开展,我们将该框架与通用VLM、特定任务预测器以及权重适配方法进行对比,以评估其专业性能、统一报告能力和适配效率。在数值、分类及报告生成指标上,所提框架大幅提升了冻结VLM的直接推理性能,在所有评估方法中实现了最强的整体性能;其新增的可训练参数仅为冻结VLM参数总量的0.006%。当排名第一的检索案例携带正确目标类别时,该框架修正了权重适配基线所犯的70.5%的错误。这些发现表明,上下文融合框架是一种轻量且有效的冻结VLM专用化策略。
英文摘要
Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record. General-purpose vision-language models (VLMs) offer a unified interface for image understanding and report generation. Existing specialization strategies, however, typically rely on task-specific models or model-weight adaptation, leaving unresolved how to introduce reliable specialist knowledge while preserving both this unified interface and the VLM's pretrained capabilities. We introduce a context-fusion framework that specializes a frozen general-purpose VLM through both implicit instruction context and explicit transduction context without modifying its pretrained weights. Specifically, a self-supervised polyp encoder retrieves related image-report pairs as explicit, query-specific evidence, while learned continuous specialist tokens provide implicit instruction context shared across cases. Experiments were conducted on 2,056 expert-annotated public endoscopic images. We compared the framework with general-purpose VLMs, task-specific predictors, and weight-adaptation methods to assess specialist performance, unified reporting, and adaptation efficiency. Across numerical, categorical, and report-generation metrics, the proposed framework substantially improved direct frozen-VLM inference and achieved the strongest overall performance among the evaluated methods. It added trainable parameters equal to only 0.006% of the frozen VLM's parameter count. When the top-1 retrieved case carried the correct target category, our framework corrected 70.5% of the errors made by a weight-adaptation baseline. These findings support the context-fusion framework as a lightweight and effective strategy for specialist adaptation of a frozen VLM.