CARE:面向医学大语言模型的因果对齐推理探索
CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models
- University of Macau(澳门大学)
- Auckland University of Technology(奥克兰理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对医学大语言模型的推理缺陷,提出CARE框架,通过因果充分性与近端可学习性条件筛选训练经验,在医学基准上提升了推理一致性与训练稳定性。
AI中文摘要:
大型语言模型(LLMs)在医学推理领域展现出强大潜力,但专家标注数据的稀缺性与高昂成本制约了其发展。尽管强化学习提供了可扩展的替代方案,医学领域标准的基于结果的方法常出现自回归信用分配失效和梯度方差爆炸问题,导致模型陷入“答案正确、推理错误”的陷阱,无意中强化虚假关联与数据集捷径,而非有效的临床推演。本研究提出Causally-Aligned Reasoning Exploration(CARE),一种基于理论的内在经验筛选框架。CARE建立在高质量训练轨迹的两个严格条件上:一是因果充分性,它利用基于一致性的自验证机制模拟do-演算干预,有效实现梯度去偏;二是近端可学习性,它采用动态熵界选择模型近端发展区内的经验,用于方差受限优化。这些经严格筛选的经验通过双流目标函数优化,该函数结合了在线策略组相对探索与难度加权经验回放。在多种医学多模态及纯文本基准上开展的大量实验表明,CARE始终优于其他强劲竞争对手,大幅减少了“答案正确但推理不一致”的情况,并提升了训练稳定性。
英文摘要:
Large Language Models (LLMs) have shown strong potential for medical reasoning, yet the scarcity and cost of expert-annotated data constrain their progress. While reinforcement learning offers a scalable alternative, standard outcome-based methods in medicine often suffer from autoregressive credit assignment failure and gradient variance explosion. This leads to the "Right Answer, Wrong Reason" trap, where models inadvertently reinforce spurious correlations and dataset shortcuts rather than valid clinical deduction. In this work, we propose Causally-Aligned Reasoning Exploration (CARE), a theoretically grounded framework for intrinsic experience curation. CARE is built upon two rigorous conditions for high-quality training trajectories: Causal Sufficiency, which utilizes an agreement-based self-verification mechanism to mimic $do$-calculus interventions and effectively debias gradients; and Proximal Learnability, which employs dynamic entropy bounds to select experiences within the model's zone of proximal development for variance-bounded optimization. These rigorously filtered experiences are optimized via a dual-stream objective that combines on-policy group-relative exploration with difficulty-weighted experience replay. Extensive experiments on diverse medical multimodal and text-only benchmarks demonstrate that CARE consistently outperforms other strong competitors, substantially reducing correct-but-inconsistent reasoning and improving training stability.