CARE:用于可靠医学视觉问答的置信感知推理
CARE: Confidence-Aware Reasoning for Reliable Medical VQA
浏览论文内容
中文总结 AI 辅助
该研究针对医学多模态大语言模型的置信度校准问题,提出CARE框架,通过双阶段流程优化准确率与校准度,在三个医学VQA基准上取得最优性能,为临床决策支持提供可信基础。
中文摘要 AI 辅助
强化微调(RFT)已使医学多模态大语言模型(MLLMs)能够为视觉问答生成思维链(CoT)推理,但这些模型存在置信度校准错误——即表达的确定性与实际诊断准确率之间的系统性差距,这会损害临床信任。我们提出CARE,一个置信感知医学推理框架,通过双阶段流程联合优化准确率与校准度:第一阶段,可扩展的医学思维链(Medical-CoT)合成为监督微调提供结构化冷启动数据;第二阶段,带新型置信感知奖励(CAR)机制的组相对策略优化(GRPO)在奖励信号中将模型的置信度与诊断正确性关联起来。在三个医学视觉问答基准上,CARE实现了最高诊断准确率,同时获得最低预期校准误差与幻觉率,为可信赖的临床决策支持奠定了基础。我们的代码可在此https URL获取。
英文摘要
Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these models suffer from $\textit{confidence miscalibration}$---a systematic gap between expressed certainty and actual diagnostic accuracy that undermines clinical trust. We propose $\textbf{CARE}$, a $\textbf{C}$onfidence-$\textbf{A}$ware medical $\textbf{RE}$asoning framework that jointly optimizes accuracy and calibration through a dual-stage pipeline. First, a scalable Medical-CoT synthesis provides structured cold-start data for Supervised Fine-Tuning. Second, Group Relative Policy Optimization (GRPO) with a novel $\textbf{Confidence-Aware Reward (CAR)}$ mechanism ties the model's confidence to diagnostic correctness within the reward signal. Across three Medical VQA benchmarks, $\textbf{CARE}$ achieves the highest diagnostic accuracy while obtaining the lowest Expected Calibration Error and Hallucination Rate, establishing a foundation for trustworthy clinical decision support. Our code is available at https://github.com/anotherbricki/CARE.
发表机构
- Ant Group(蚂蚁集团)
- University of Michigan(密歇根大学)
- City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。