arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10996cs.CL

ConRub-Med:面向开放式医学问答的共识准则强化学习

ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering

Taojie Zhu, Yuan Xia, Tao Sun, Yizhi Wang, Yan Chen, Qunshan He, Tian Guan, Jian Wang, Jinjie Gu, Junwei Liu, Yonghong He

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出ConRub-Med,通过三模型共识准则与三状态评分优化GRPO策略,在9个医学问答基准中6个排第一,HealthBench-Hard得分优于InfiMed-ORBIT。

中文摘要 AI 辅助

可验证奖励的强化学习在数学和编码领域成效显著,因其答案可自动校验。但许多开放式医学问题缺乏同等低成本的结果校验工具:回答可能部分正确、不完整或包含临床相关错误。由医生撰写或验证的准则提供了坚实的临床依据,但让专家参与每个实例的成本高昂。模型生成的准则使这种监督具有可扩展性。我们提出ConRub-Med,以在准则反馈从构建到策略优化的过程中保留有用的区分度。对于每个提示,三个异构语言模型独立提出原子准则;一个单独的模型对这些准则进行审核,仅保留获得所有三个生成器语义支持的准则。三状态评分区分正确覆盖、缺失信息和错误主张,错误会获得负分而非零分。当完整的分组相对策略优化(GRPO)组中的每个回答获得相同的最终奖励时,只有当两个候选排序一致时,成对评判者才会提供序列优势,且不改变标量奖励;无平局的组使用普通GRPO。在一项按问题匹配的盲法研究中,两名医学专家对完整流程生成的回答面板的临床相关性评分高于单个生成器生成的面板。在评估的开放模型中,ConRub-Med在9个基准中的6个上排名第一,且在医学和泛化平均性能上达到最高。利用生成的包含5166个提示的准则数据集,其在HealthBench-Hard上的得分为38.98±1.04(均值±标准差),而InfiMed-ORBIT使用8000个样本的得分为33.60,使用28000个样本的得分为37.30。

英文摘要

Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many open-ended medical questions lack comparably cheap outcome verifiers: responses may be partly correct, incomplete, or contain clinically consequential errors. Rubrics written or validated by physicians offer strong clinical grounding, but involving experts in every instance is costly. Model-generated rubrics make this supervision scalable. We introduce ConRub-Med to preserve useful distinctions as rubric feedback moves from construction to policy optimization. For each prompt, three heterogeneous language models propose atomic criteria independently; a separate model reviews them, retaining only criteria with semantic support from all three generators. Three-State scoring distinguishes correct coverage, missing information, and incorrect claims. Errors receive negative rather than zero credit. When every response in a complete Group Relative Policy Optimization (GRPO) group receives the same final reward, a pairwise judge provides sequence advantages only if both candidate orders agree, without changing the scalar rewards. Groups without ties use vanilla GRPO. In a blinded study matched by question, two medical experts rate panels from the full pipeline as more clinically relevant than panels produced by one generator. Across the evaluated open models, ConRub-Med ranks first on six of nine benchmarks and achieves the highest medical and generalization averages. Using the resulting rubric dataset of 5,166 prompts, it scores $38.98 \pm 1.04$ (mean $\pm$ SD) on HealthBench-Hard, compared with InfiMed-ORBIT's 33.60 with 8,000 samples and 37.30 with 28,000.

↑