arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08623cs.AI

MedCalc-R1:面向医学数学推理的知识引导奖励框架

MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning

  • Harbin Institute of Technology(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Haotian Wang, Lian Yan, Xingzhi Yao, Fanshu Meng, Ye He, Jingchi Jiang, Yi Guan

AI总结:

该研究针对医学数学推理中RLVR框架的奖励评估缺陷,提出知识引导的混合奖励框架MedCalc-R1,通过强制公式生成与验证、软硬奖励方案提升性能,在安全关键领域表现更优。

AI中文摘要:

在用于数学推理任务的带可验证奖励的强化学习(RLVR)框架中,浮点结果通常采用基于容差的奖励进行评估。然而,该策略存在阈值校准困难、训练动态不稳定、精度有限等挑战,尤其在临床场景中问题突出。为解决这些局限,我们提出知识引导的混合奖励框架(\textsc{MedCalc-R1})。具体而言,我们引入知识验证奖励机制,强制模型显式生成计算公式,并通过外部验证器对公式进行验证,以提升可解释性和推理可靠性。此外,我们设计混合软硬奖励方案,将基于临床安全阈值的硬约束与精度敏感的软奖励相结合,在可接受范围内逐步引导学习。实验结果表明,我们的方法在推理精度和泛化能力上均显著优于现有基线,验证了其在安全关键领域的有效性和适用性。

英文摘要:

In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically evaluated using a tolerance-based reward. However, this strategy suffers from challenges such as difficulty in threshold calibration, unstable training dynamics, and limited accuracy, especially in clinical scenarios. To address these limitations, we propose a knowledge-guided hybrid reward framework (\textsc{MedCalc-R1}). Specifically, we introduce a knowledge verification reward mechanism that enforces explicit generation of computational formulas, which are further validated by an external verifier to enhance interpretability and reasoning reliability. Furthermore, we design a hybrid soft-hard reward scheme combining a hard constraint based on clinical safety thresholds with a soft, precision-sensitive reward that progressively guides learning within the acceptable range. Experimental results demonstrate that our method significantly outperforms existing baselines in both reasoning accuracy and generalization capability, validating the effectiveness and applicability in safety-critical domains.

补充信息

↑