分离能力与置信度:面向GRPO训练的医学视觉语言模型的接地双状态校准
Separating Capability from Confidence: Grounded Dual-State Calibration for GRPO-Trained Medical Vision-Language Models
- Shanghai Jiao Tong University(上海交通大学)
- Medical Image Insights Co. Ltd.(医学影像洞察有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对GRPO训练的医学视觉语言模型,提出DualRead方法,分离能力学习与置信度估计,结合回答前后状态校准,并引入CCG-AUC评估视觉接地,提升置信度判别与校准。
AI中文摘要:
医学视觉语言模型(VLM)需要反映答案正确性和患者特定视觉证据的置信度。近期基于GRPO的方法将口头化置信度与答案生成联合优化。然而,这种联合优化可能干扰答案学习,并使置信度趋向于接近二值的数值。口头化置信度也无法提供对视觉支持的明确评估。因此,我们将能力学习与置信度估计分离,并提出\textbf{DualRead}。DualRead基于一个洞察:可靠性可以从演员在回答过程关键时刻的内部状态中读取。它冻结GRPO训练的演员,并将回答前的可解性与回答后对生成答案及其视觉支持的评估相结合。为进一步评估置信度是否反映视觉接地,我们引入\textbf{反事实置信度接地AUC}(CCG-AUC)。它衡量当真实图像替换使演员从正确变为错误时,置信度是否下降。在两个VLM骨干网络以及分布内和分布外的医学视觉问答基准上,DualRead在保持答案准确性的同时,相比口头化置信度提高了正确性判别和校准。CCG-AUC揭示了置信度是否响应与答案相关的视觉证据,而非主要响应非视觉线索。
英文摘要:
Medical vision-language models (VLMs) require confidence that reflects both answer correctness and patient-specific visual evidence. Recent GRPO-based methods optimize verbalized confidence together with answer generation. However, this joint optimization may interfere with answer learning and drive confidence toward near-binary values. Verbalized confidence also provides no explicit assessment of visual support. We therefore separate capability learning from confidence estimation and propose \textbf{DualRead}. DualRead builds on the insight that reliability can be read from the actor's internal states at critical moments in the answering process. It freezes the GRPO-trained actor and combines pre-answer solvability with a post-answer assessment of the generated answer and its visual support. To further assess whether confidence reflects visual grounding, we introduce \textbf{Counterfactual Confidence Grounding AUC} (CCG-AUC). It measures whether confidence decreases when real-image substitution changes the actor from correct to incorrect. Across two VLM backbones and both in- and out-of-distribution medical VQA benchmarks, DualRead improves correctness discrimination and calibration over verbalized confidence while preserving answer accuracy. CCG-AUC reveals whether confidence responds to answer-relevant visual evidence rather than primarily to non-visual cues.