发表机构
The University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出轻量级探测框架MedProb,基于冻结VLM表征预测多项选择式Med-VQA答案,在多个医学VQA数据集上性能优于提示方法、医学VLMs及智能体系统,还可扩展至开放式生成。
AI 中文摘要
医学视觉问答(Med-VQA)常被认为需要医学微调、大模型或复杂多智能体流水线。本文重新审视这一假设,提出MedProb——一种轻量级探测框架,可从冻结的视觉语言模型(VLM)表征中预测多项选择式Med-VQA答案,无需自由文本生成。在PATH-VQA、SLAKE和VQA-RAD数据集上,MedProb恢复的答案相关信号远多于提示方法,性能优于医学VLMs和智能体系统。与提示相比,探测缩小了大小模型间的表观差距,表明小型VLMs所含可恢复的Med-VQA信号多于基于生成的评估所揭示的内容。在14组匹配的通用与医学VLM对中,医学适配并未始终提升这种线性可解码性。最后,自由文本生成存在最高达10个百分点的答案位置偏差,而MedProb虽也存在位置偏差,但受影响的方式与提示不同。本文主要结果针对多项选择/多类Med-VQA场景;还展示了该探测可通过拒绝采样评分程序扩展至开放式生成。
英文摘要
Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or complex multi-agent pipelines. We revisit this assumption with \textbf{MedProb}, a lightweight probing framework that predicts multiple-choice Med-VQA answers from frozen VLM representations without free-text generation. Across PATH-VQA, SLAKE, and VQA-RAD, MedProb recovers substantially more answer-relevant signal than prompting and performs stronger than medical VLMs and agentic systems. Probing also reduces the apparent gap between small and large models compared to prompting, suggesting that smaller VLMs contain more recoverable Med-VQA signal than generation-based evaluation reveals. Across 14 matched general-purpose and medical VLM pairs, medical adaptation does not consistently improve this linear decodability. Finally, free-text generation exhibits an answer-position bias of up to 10 percentage points, whereas MedProb also has positional bias, however, it is impacted differently than prompting. Our main results target the multiple-choice/multiclass Med-VQA setting; we additionally show the probe can be extended to open-ended generation via a rejection-sampling scoring procedure.
CommentsAccepted to EMNLP Findings 2026