发表机构
Institute of Business and Administration(商业管理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对胃肠道内窥镜视觉问答任务,提出多任务微调方案,利用现有数据集构建辅助任务,微调小型视觉语言模型,实现了准确率提升与答案-图像区域对齐优化。
AI 中文摘要
胃肠道(GI)内窥镜图像分析已从单标签分类转向视觉问答(VQA),要求模型回答关于图像的自由格式临床问题。尽管近期视觉语言模型(VLMs)在该任务上取得了可观的回答准确率,但临床应用还需模型的内部表征能反映其答案背后的视觉证据。我们提出一种简单的多任务微调方案,利用现有VQA数据集构建辅助的 grounding 和描述任务,仅需极少额外标注:专家标注的息肉掩码可直接复用,而带有Grad-CAM定位的GI域预训练分类器为缺乏真实掩码的类别提供弱监督。我们在Kvasir-VQA-x1数据集上,针对仅VQA和多任务两种设置,用低秩适配微调了三种小型VLM主干,结果显示在分布内和分布外数据上,模型不仅取得了持续的准确率提升,还改进了答案 token 与相关图像区域的隐式对齐。
英文摘要
Gastrointestinal (GI) endoscopic image analysis has shifted from single-label classification toward visual question answering (VQA), where a model must answer free-form clinical questions about an image. While recent vision-language models (VLMs) achieve promising answer accuracy on this task, clinical adoption also requires the model's internal representations to reflect the visual evidence behind its answers. We propose a simple multi-task fine-tuning recipe that constructs auxiliary grounding and description tasks from an existing VQA dataset with minimal additional annotation: expert-annotated polyp masks are reused directly, while a GI-domain pretrained classifier with Grad-CAM localization provides weak supervision for finding categories that lack ground-truth masks. Three small VLM backbones are fine-tuned with low-rank adaptation under matched VQA-only and multi-task recipes on Kvasir-VQA-x1, and we show consistent accuracy gains together with improved implicit alignment between answer tokens and the relevant image region, evaluated on both in-distribution and out-of-distribution data.
CommentsAccepted at EMA4MICCAI 2026 (Workshop on Efficient Medical AI, MICCAI 2026)