arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从奖励信号到视觉效用:医学视觉语言模型后训练的控制性审计

From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training

Wang Jingxin

arXiv 2609.31450首次发表:更新:

发表机构

Institute of Neuroscience, CAS(中国科学院神经科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控实验审计医学视觉语言模型后训练,发现仅语言模型LoRA SFT在准确率提升有限但损害视觉效用,而扩展适应范围和标准GRPO效果不佳,证据目标在训练集有改进但验证集不稳定,揭示了优化活动与有用视觉行为之间的差距。

AI 中文摘要

医学视觉语言模型(VLM)的后训练通常通过答案准确率来评估。我们在一项受控的Qwen2.5-VL-3B研究(基于PMC-VQA数据集)中,考察了准确率变化和训练目标如何与基于图像的决策相关联。我们比较了仅限语言模型的低秩适应(LoRA)监督微调(SFT)、扩展的多模态适应范围、标准的仅答案组相对策略优化(GRPO),以及一种反事实证据目标。在2000个干净测试问题上,语言模型LoRA SFT使正确图像准确率变化了+1.10个百分点(95%配对自助法置信区间:-0.85至+3.05),而视觉效益事件减少了2.40个百分点,图像敏感性下降了5.60个百分点。配对记录显示,有155个获得和203个丢失的视觉效益事件。更广泛的适应范围产生的正确图像准确率低于语言模型LoRA SFT。标准GRPO产生了混合奖励组和参数更新,其干净测试准确率变化不确定。生成审计揭示,规范选项分数可能遵循与生成答案不同的token路径。当沿着贪婪生成路径取分数时,证据目标在训练集上有所改进;在匹配的训练剂量下,其在验证数据上相对于标准GRPO的优势仍不一致。样本级分析追踪了后训练期间证据分数、决策边界和生成答案的变化。这项实证和测量审计识别了优化活动、目标获取和有用的保留视觉行为之间的差距。

英文摘要

Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy. We examine how changes in accuracy and training objectives relate to image-conditioned decisions in a controlled Qwen2.5-VL-3B study on PMC-VQA. We compare supervised fine-tuning (SFT) with low-rank adaptation (LoRA) restricted to the language model, expanded multimodal adaptation scopes, standard answer-only Group Relative Policy Optimization (GRPO), and a counterfactual evidence objective. On 2,000 clean-test questions, language model LoRA SFT changes correct-image accuracy by +1.10 percentage points (95% paired bootstrap CI:-0.85 to +3.05), while visual-benefit events decrease by 2.40 points and image sensitivity decreases by 5.60 points. Paired records reveal 155 acquired and 203 lost visual-benefit events. Broader adaptation yields lower correct-image accuracy than language-model LoRA SFT. Standard GRPO produces mixed-reward groups and parameter updates, with an uncertain clean test accuracy change. A generation audit reveals that canonical option scores can follow a different token path from generated answers. With scores taken along the greedy generation path, the evidence target improves on the training set; its gains over standard GRPO remain inconsistent on validation data at matched training doses. Sample-level analyses trace how evidence scores, decision margins, and generated answers change during post-training. This empirical and measurement audit identifies gaps between optimization activity, target acquisition, and useful held-out visual behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑