arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

冻结模型,进化专长:面向多模态医学AI的部署经验模型无关学习

Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI

Yexiao He, Yucheng Tang, Pengfei Guo, Yufan He, Andriy Myronenko, Can Zhao, Ang Li, Daguang Xu, Dong Yang

arXiv 2610.09146首次发表:更新:

发表机构

University of Maryland, College Park; NVIDIA(马里兰大学学院公园分校; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出模型无关框架,通过技能、知识记忆和多模态知识库三种外部经验,使冻结的医学多模态模型在线部署中持续学习,在六个基准上性能提升高达34.2%。

AI 中文摘要

大型语言模型(LLMs)和视觉-语言模型(VLMs)在部署后通常被冻结,因此它们无法从所解决的病例中学习。这在医学领域尤其令人担忧,因为新的临床证据、更新的指南和新的疗法可能改变既定实践。微调可以更新模型,但需要访问模型权重并进行额外训练。无参数方法避免了训练,但它们可能过拟合固定的验证集,缺乏可靠的领域知识,或因仅以文本形式保存经验而丢失视觉细节。为解决这些局限性,我们提出一个模型无关框架,允许冻结的LLMs和VLMs通过三种形式的外部专业知识从部署经验中学习:一个指导推理和工具使用的技能(Skill),一个存储由先前病例或可信外部证据支持的可信事实的知识记忆(Knowledge Memory),以及一个保存视觉示例并引导模型将每个检索到的病例与当前图像相关联的多模态知识库(Multimodal Knowledge Base)。不依赖固定验证集,验证策略仅当更新有助于新病例且不降低先前病例性能时才保留更新。在覆盖临床诊断、临床工作流程、医学推理以及医学和非医学视觉推理的六个基准上,并使用四个开放权重和闭源基础模型,我们的框架在在线部署期间将医学任务上的性能比基础模型提升高达34.2%,泛化到未见病例,无需进一步优化即可迁移到其他模型,并在非医学领域同样有效。

英文摘要

Large language models (LLMs) and vision-language models (VLMs) are usually frozen after deployment, so they do not learn from the cases they solve. This is especially concerning in medicine, where new clinical evidence, updated guidelines, and new therapies can change established practice. Fine-tuning can update the model, but it requires access to model weights and additional training. Parameter-free methods avoid training, but they may overfit a fixed validation set, lack reliable domain knowledge, or lose visual details by saving experience only as text. To address these limitations, we present a model-agnostic framework that allows frozen LLMs and VLMs to learn from deployment experience through three forms of external expertise: a Skill that guides reasoning and tool use, a Knowledge Memory that stores reliable facts supported by earlier cases or trusted external evidence, and a Multimodal Knowledge Base that keeps visual examples and guides the model to relate each retrieved case to the current image. Instead of relying on a fixed validation set, a validation strategy keeps an update only if it helps on new cases without degrading performance on earlier ones. Across six benchmarks covering clinical diagnosis, clinical workflows, medical reasoning, and medical and non-medical visual reasoning, and with four open-weight and closed-source base models, our framework improves performance during online deployment by up to 34.2% over the base model on medical tasks, generalizes to unseen cases, transfers to other models without further optimization, and works in non-medical domains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑