arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24057cs.AIcs.CLcs.CV

表示引导的上下文学习用于多模态大语言模型的医学图像解读

Representation-guided in-context learning for medical image interpretation with multimodal large language models

发表机构麻省眼耳医院谢本斯眼科研究所 · 哈佛医学院 · 麻省眼耳医院
另 6 家 · 查看机构详情
  • Schepens Eye Research Institute of Massachusetts Eye and Ear(麻省眼耳医院谢本斯眼科研究所)
  • Harvard Medical School(哈佛医学院)
  • Massachusetts Eye and Ear(麻省眼耳医院)
  • National Taiwan University Hospital(台湾大学医院)
  • Massachusetts Institute of Technology(麻省理工学院)
  • Northwestern University(西北大学)
  • Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(哈佛大学肯普纳自然与人工智能研究所)
  • Harvard University(哈佛大学)
  • Harvard Data Science Initiative, Harvard University(哈佛大学数据科学倡议)

机构由 AI 辅助整理,请以论文原文为准。

Minda Zhao, Fangyu Hu, Yan Luo, Yutong Yang, Jiahui Cai, Kaichen Zhou, Manling Li, Paul Liang, Yilun Du, Lucy Q. Shen, Mengyu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

提出表示引导的上下文学习(RG-ICL),通过冻结编码器检索对齐案例,免训练提升医学图像分类和VQA性能,表明案例选择可替代重训练。

中文摘要 AI 辅助

医学图像解读是诊断和护理的核心,但将通用多模态大语言模型(MLLMs)适应于此通常需要资源密集型的领域特定微调。在此,我们引入了表示引导的上下文学习(RG-ICL),这是一种免训练的推理框架,利用冻结的编码器检索与查询对齐的演示样本,无需任务特定的参数更新。在涵盖组织病理学、放射学和视网膜眼底检查的八个数据集上,与无上下文和传统ICL相比,RG-ICL在分类(平均提升20个百分点)和视觉问答(VQA)(平均提升13个百分点)方面均有改进,接近或超过了基于训练的比较方法。检索哪些案例比检索多少更重要:6个查询对齐的案例优于多达32个随机选择的案例,而固定或随机的案例往往将准确率降至基线以下。对于VQA,将参考案例与图像内容和问题意图对齐产生了进一步的增益。这些发现表明,对于医学图像解读,策划MLLM所见的参考案例是重新训练的一种实用替代方案。

英文摘要

Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires resource-intensive domain-specific fine-tuning. Here we introduce representation-guided in-context learning (RG-ICL), a training-free inference framework that retrieves query-aligned demonstrations using frozen encoders, without task-specific parameter updates. Across eight datasets spanning histopathology, radiology and retinal fundoscopy, RG-ICL improved classification (mean gain 20 percentage points) and visual question answering (VQA) (mean gain 13 percentage points) over no-context and conventional ICL, approaching or exceeding training-based comparators. Which cases were retrieved mattered more than how many: 6 query-aligned cases outperformed up to 32 randomly selected ones, whereas fixed or random cases often reduced accuracy below baseline. For VQA, aligning reference cases with both image content and question intent produced further gains. These findings indicate that for medical image interpretation, curating which reference cases an MLLM sees is a practical alternative to retraining it.

↑