发表机构
Tsinghua University; Meituan(清华大学; 美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对大语言模型后训练中 rubric 进化的问题,提出 CARE 方法,通过锚点响应对比实现自适应与 Chase 分支,在多个基准测试中达到最优性能且胜率持续提升,可跨模型家族泛化。
AI 中文摘要
基于 rubric 的强化学习将开放式指令分解为特定于提示、灵活的 rubric,相比使用可验证奖励的强化学习,更适合对大语言模型(LLM)进行开放式任务的后训练。然而,静态 rubric 会随着策略的进化而被破解,现有动态方法会引入新问题:无定向的 rubric 提取、不可靠的破解检测以及无约束的 rubric 增殖。我们提出 CARE(Contrastive Anchor-based Rubric Evolution,基于锚点的对比式 rubric 进化),该方法将每一步 rubric 进化都基于前沿模型根据提示及其 rubric 生成的高质量锚点响应。在每一步训练中,CARE 将得分最高的 rollout 与锚点进行对比,实现两种互补机制:自适应分支可被动修复奖励错误指定;Chase 分支可主动将前沿级别的质量差距转化为更清晰的 rubric。两个分支共同维持高奖励区域的判别准确率,而该区域正是奖励过度优化的主要来源。在 WildChecklist-9K 数据集上使用 Qwen2.5-7B-Base 和 Qwen2.5-7B-Instruct 进行的实验表明,CARE 在 Arena-Hard-2.0、InfoBench 和 FollowBench 上实现了最先进的性能,并且是唯一一种在 300 步训练中对 GPT-4.1 锚点响应的胜率持续提升的方法;在 Llama-3.1-8B-Instruct 和 Qwen3-8B 上的额外结果进一步表明,CARE 可跨模型家族泛化。
英文摘要
Rubric-based reinforcement learning decomposes open-ended instructions into prompt-specific, flexible rubrics, making it better suited than reinforcement learning with verifiable rewards for post-training LLMs on open-ended tasks. However, static rubrics are inevitably hacked as the policy evolves, and existing dynamic approaches introduce new problems: undirected rubric extraction, unreliable hack detection, and unbounded rubric proliferation. We propose $\textbf{CARE}$ ($\textbf{C}$ontrastive $\textbf{A}$nchor-based $\textbf{R}$ubric $\textbf{E}$volution), which grounds every rubric evolution step in a high-quality anchor response generated by a frontier model conditioned on the prompt and its rubrics. At each training step, CARE contrasts the highest-scoring rollout against the anchor, enabling two complementary mechanisms: an Adaptive branch that reactively repairs reward misspecification; and a Chase branch that proactively converts frontier-level quality gaps into sharper rubrics. Together, the two branches $\textbf{maintain discriminative accuracy in the high-reward region}$---the precise region where reward over-optimization mostly originates. Experiments on WildChecklist-9K with Qwen2.5-7B-Base and Qwen2.5-7B-Instruct show that CARE achieves state-of-the-art performance on Arena-Hard-2.0, InfoBench, and FollowBench, and is the $\textbf{only}$ method whose win rate against GPT-4.1 anchor responses shows sustained improvement throughout 300 training steps; additional results on Llama-3.1-8B-Instruct and Qwen3-8B further indicate that CARE generalizes across model families.
CommentsEMNLP 2026 MainConference