发表机构
NTU Singapore(新加坡南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MetaRubric通过证据感知策略优化与响应引导的评分标准适应,解决基于评分标准的强化学习中的空洞信用问题,显著提升医学问答等任务的准确率。
AI 中文摘要
基于评分标准的强化学习通过为各个响应要求分配部分分数,将奖励驱动的优化扩展到开放式任务。然而,评分标准评判者即使响应中缺少其所需的信息或操作,也可能给出较高的标准分数,我们将这种失败模式称为“空洞信用”。这种奖励在所需信息被移除后仍然存在,并且可能逆转响应GRPO优势的符号。为解决此问题,我们引入了MetaRubric,该方法交替进行证据感知的策略优化与响应引导的评分标准适应。我们通过在每个提示中更改一个与任务相关的事实来构建反事实对应项。在策略优化期间,仅当响应包含足够的证据以满足所需评分标准时,才分配信用。在每个策略优化阶段之后,当前策略的响应会指导对原始和反事实标准的修订,同时保留原始提示初始评分标准在每个提示事实下所解释的含义。我们还在阶段边界调整标准权重,以更好地解决观察到的策略错误。在多个骨干网络上,MetaRubric在PubMedQA上的准确率比静态评判者GRPO提高了6.00至20.40个百分点,并在HealthBench-Hard和两个多模态医学基准上取得了进一步的提升。
英文摘要
Rubric-based reinforcement learning extends reward-driven optimization to open-ended tasks by assigning partial credit to individual response requirements. However, rubric judges can assign a high criterion score even when the information or action it requires is absent from the response, a failure mode we term Vacuous Credit. Such awards persist after the required information is removed and can reverse the sign of a response's GRPO advantage. To address this problem, we introduce MetaRubric, which alternates evidence-aware policy optimization with response-guided rubric adaptation. We construct counterfactual counterparts by changing one task-relevant fact in each prompt. During policy optimization, credit is assigned only when the response contains sufficient evidence to satisfy the required rubric criterion. After each policy-optimization stage, current policy responses guide revisions to original and counterfactual criteria while preserving the meaning of the original prompt's initial rubric as interpreted under each prompt's facts. We also adapt criterion weights at stage boundaries to better address observed policy errors. Across multiple backbones, MetaRubric improves PubMedQA accuracy by 6.00--20.40 percentage points over static-judge GRPO, with further gains on HealthBench-Hard and two multimodal medical benchmarks.
CommentsProject page:https://metarubric.github.io/