arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CARE:基于锚点的对比式 rubric 进化用于大语言模型后训练

CARE: Contrastive Anchor-based Rubric Evolution for Large Language Model Post-Training

Siyuan Li, Xinxin Song, Chen Ruinian, Jingjing Fan, Tingxiong Xiao, Yangen Hu, Ke Zeng, Jinli Suo

arXiv 2609.00892首次发表:更新:

发表机构

Tsinghua University; Meituan(清华大学; 美团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对大语言模型后训练中 rubric 进化的问题,提出 CARE 方法,通过锚点响应对比实现自适应与 Chase 分支,在多个基准测试中达到最优性能且胜率持续提升,可跨模型家族泛化。

AI 中文摘要

基于 rubric 的强化学习将开放式指令分解为特定于提示、灵活的 rubric,相比使用可验证奖励的强化学习,更适合对大语言模型(LLM)进行开放式任务的后训练。然而,静态 rubric 会随着策略的进化而被破解,现有动态方法会引入新问题:无定向的 rubric 提取、不可靠的破解检测以及无约束的 rubric 增殖。我们提出 CARE(Contrastive Anchor-based Rubric Evolution,基于锚点的对比式 rubric 进化),该方法将每一步 rubric 进化都基于前沿模型根据提示及其 rubric 生成的高质量锚点响应。在每一步训练中,CARE 将得分最高的 rollout 与锚点进行对比,实现两种互补机制:自适应分支可被动修复奖励错误指定;Chase 分支可主动将前沿级别的质量差距转化为更清晰的 rubric。两个分支共同维持高奖励区域的判别准确率,而该区域正是奖励过度优化的主要来源。在 WildChecklist-9K 数据集上使用 Qwen2.5-7B-Base 和 Qwen2.5-7B-Instruct 进行的实验表明,CARE 在 Arena-Hard-2.0、InfoBench 和 FollowBench 上实现了最先进的性能,并且是唯一一种在 300 步训练中对 GPT-4.1 锚点响应的胜率持续提升的方法;在 Llama-3.1-8B-Instruct 和 Qwen3-8B 上的额外结果进一步表明,CARE 可跨模型家族泛化。

英文摘要

Rubric-based reinforcement learning decomposes open-ended instructions into prompt-specific, flexible rubrics, making it better suited than reinforcement learning with verifiable rewards for post-training LLMs on open-ended tasks. However, static rubrics are inevitably hacked as the policy evolves, and existing dynamic approaches introduce new problems: undirected rubric extraction, unreliable hack detection, and unbounded rubric proliferation. We propose $\textbf{CARE}$ ($\textbf{C}$ontrastive $\textbf{A}$nchor-based $\textbf{R}$ubric $\textbf{E}$volution), which grounds every rubric evolution step in a high-quality anchor response generated by a frontier model conditioned on the prompt and its rubrics. At each training step, CARE contrasts the highest-scoring rollout against the anchor, enabling two complementary mechanisms: an Adaptive branch that reactively repairs reward misspecification; and a Chase branch that proactively converts frontier-level quality gaps into sharper rubrics. Together, the two branches $\textbf{maintain discriminative accuracy in the high-reward region}$---the precise region where reward over-optimization mostly originates. Experiments on WildChecklist-9K with Qwen2.5-7B-Base and Qwen2.5-7B-Instruct show that CARE achieves state-of-the-art performance on Arena-Hard-2.0, InfoBench, and FollowBench, and is the $\textbf{only}$ method whose win rate against GPT-4.1 anchor responses shows sustained improvement throughout 300 training steps; additional results on Llama-3.1-8B-Instruct and Qwen3-8B further indicate that CARE generalizes across model families.

CommentsEMNLP 2026 MainConference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑