arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10964cs.CVcs.AI

CARE:用于可靠医学视觉问答的置信感知推理

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

Yuetian Du, Yucheng Wang, Zhenyuan Chen, Luyuan Chen, Rongyu Zhang, Jinjian Zhang, Wei Zhou, Zhijie Xu, Ming Kong, Zhan Zhou, Jie Liu, Qiang Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对医学多模态大语言模型的置信度校准问题,提出CARE框架,通过双阶段流程优化准确率与校准度,在三个医学VQA基准上取得最优性能,为临床决策支持提供可信基础。

中文摘要 AI 辅助

强化微调(RFT)已使医学多模态大语言模型(MLLMs)能够为视觉问答生成思维链(CoT)推理,但这些模型存在置信度校准错误——即表达的确定性与实际诊断准确率之间的系统性差距,这会损害临床信任。我们提出CARE,一个置信感知医学推理框架,通过双阶段流程联合优化准确率与校准度:第一阶段,可扩展的医学思维链(Medical-CoT)合成为监督微调提供结构化冷启动数据;第二阶段,带新型置信感知奖励(CAR)机制的组相对策略优化(GRPO)在奖励信号中将模型的置信度与诊断正确性关联起来。在三个医学视觉问答基准上,CARE实现了最高诊断准确率,同时获得最低预期校准误差与幻觉率,为可信赖的临床决策支持奠定了基础。我们的代码可在此https URL获取。

英文摘要

Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these models suffer from $\textit{confidence miscalibration}$---a systematic gap between expressed certainty and actual diagnostic accuracy that undermines clinical trust. We propose $\textbf{CARE}$, a $\textbf{C}$onfidence-$\textbf{A}$ware medical $\textbf{RE}$asoning framework that jointly optimizes accuracy and calibration through a dual-stage pipeline. First, a scalable Medical-CoT synthesis provides structured cold-start data for Supervised Fine-Tuning. Second, Group Relative Policy Optimization (GRPO) with a novel $\textbf{Confidence-Aware Reward (CAR)}$ mechanism ties the model's confidence to diagnostic correctness within the reward signal. Across three Medical VQA benchmarks, $\textbf{CARE}$ achieves the highest diagnostic accuracy while obtaining the lowest Expected Calibration Error and Hallucination Rate, establishing a foundation for trustworthy clinical decision support. Our code is available at https://github.com/anotherbricki/CARE.

发表机构

  • Ant Group(蚂蚁集团)
  • University of Michigan(密歇根大学)
  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑