arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16303cs.CVcs.AI

Med-OPD:通过证据感知策略蒸馏改进医学视觉语言模型

Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation

Yunhang Qian, Jiaquan Yu, Jiawei Liu, Meng Wang, Hongwei Bran Li, Xiaobin Hu

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对医学视觉语言模型依赖语言先验而非视觉证据推理的问题,提出Med-OPD框架,引入医学证据优势信号,在令牌和轨迹级别重新分配蒸馏信号,实验证明该方法能加强模型对关键视觉证据的依赖,提升多模态医学推理能力。

中文摘要 AI 辅助

医学视觉语言模型(Med-VLMs)需要从细粒度视觉证据进行可靠推理,但现有模型常依赖语言先验或医学模板给出看似合理的临床答案,而非真正关注关键诊断区域。策略蒸馏(OPD)能对学生生成轨迹进行密集令牌级监督且隐私兼容。然而,标准OPD均匀蒸馏所有令牌,使依赖证据的稀疏令牌被大量临床叙述令牌稀释。受OPD在大语言模型社区成功的启发,我们提出Med-OPD,这是首个将策略蒸馏与医学证据感知监督集成的统一训练后框架。我们引入医学证据优势(MEA),一种基于教师的反事实信号,通过答案感知提示聚焦教师对支持目标诊断证据的评分,并通过比较原始和证据退化成像模式下教师的可能性来衡量每个令牌对医学视觉证据的依赖。基于MEA,Med-OPD在令牌和轨迹级别重新分配蒸馏信号,强调关键诊断令牌和依赖证据的展开。在OmniMedVQA子集上的实验表明,Med-OPD在CT、MRI、疾病诊断和病变分级方面始终优于SFT和标准OPD。这些结果表明,证据感知蒸馏可以更好地加强医学VLMs对关键视觉证据的依赖,提高可靠的多模态医学推理能力。源代码和数据可在指定链接公开获取。

英文摘要

Medical Vision-Language Models (Med-VLMs) require reliable reasoning from fine-grained visual evidence, yet existing models can produce plausible clinical answers by relying on language priors or medical templates rather than truly attending to diagnosis-critical regions. On-Policy Distillation (OPD) offers dense token-level supervision on student-generated trajectories and provides a privacy-compatible means of capability transfer without requiring the redistribution of raw patient data. However, standard OPD uniformly distills all tokens, causing sparse evidence-dependent tokens to be diluted by abundant clinical narrative tokens. Inspired by the success of OPD in the large language model community, we propose \textbf{Med-OPD}, to our knowledge the first unified post-training framework that integrates on-policy distillation with medical evidence-aware supervision for Med-VLMs. We introduce \textbf{Medical Evidence Advantage} (MEA), a teacher-grounded counterfactual signal that uses an answer-aware hint to focus teacher scoring on evidence supporting the target diagnosis, and measures each token's dependence on medical visual evidence by comparing teacher likelihoods under the original and evidence-degraded imaging modalities. Based on MEA, Med-OPD redistributes the distillation signal at both the token and trajectory levels, emphasizing diagnosis-critical tokens and evidence-reliant rollouts. Experiments on OmniMedVQA subsets show that Med-OPD consistently outperforms SFT and standard OPD across CT, MRI, Disease Diagnosis, and Lesion Grading. These results demonstrate that evidence-aware distillation can better strengthen medical VLMs' reliance on key visual evidence and improve reliable multimodal medical reasoning. The source code and data is publicly available at: https://github.com/yunhang8658/MedOPD.git

发表机构

  • National University of Singapore(新加坡国立大学)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑