作为特权信息的评分规则:面向开放式生成
Rubrics as Privileged Information for Open-Ended Generation
浏览论文内容
中文总结 AI 辅助
该研究将在线策略自蒸馏扩展至开放式生成,提出以评分规则作为特权信息(RuPI),在Qwen、Llama等模型上,其性能优于评分规则作为奖励的RL及参考完成PI蒸馏方法。
中文摘要 AI 辅助
在线策略自蒸馏(OPSD)是指单个模型在不同上下文下同时充当学生模型和教师模型,该方法在数学等可验证领域已展现出应用前景,这类领域中存在以真实答案形式呈现的硬特权信息(PI),可从结构层面约束有效续答。我们将OPSD扩展至开放式生成任务,采用以评分规则(rubrics)形式呈现的软PI,这类规则用于引导偏好但允许多种有效响应。评分规则已作为强化学习(RL)中的标量奖励使用;我们证明,在蒸馏任务中,它们能提供丰富得多的信号作为密集PI,且与直觉相反,在该场景下,软评分规则PI在学生模型的输出上提供的训练信号比硬参考完成PI更丰富、更有效。参考完成是一组有效响应中的一个点,因此向其蒸馏会过度约束学生模型,而评分规则则指定了一组有效响应共有的偏好结构。我们在Qwen和Llama模型系列上展示了将评分规则作为PI用于开放式生成的有效性,且其性能优于采用评分规则作为奖励(RaR)的RL方法,实验使用HealthBench基准,该基准根据医生创建的评分规则对开放式健康响应进行评分,为开放式任务提供密集的 token 级监督;RuPI在三个模型上相比RaR RL的绝对得分提升最高达+0.10,在匹配配方和KL方向下,相比参考PI的绝对得分提升为+0.034至+0.079。我们进一步证明这些发现可推广至在RubricHub Science语料上训练并在ResearchQA上评估的场景:软评分规则PI的表现优于参考PI蒸馏和RaR RL(66.6%对比64.2%和57.6%)。
英文摘要
On-policy self-distillation (OPSD), where a single model acts as both student and teacher with different contexts, has shown promise in verifiable domains like math, where hard privileged information (PI) in the form of ground-truth answers structurally constrains valid continuations. We extend OPSD to open-ended generation using soft PI in the form of rubrics that guide preferences but admit many valid responses. Rubrics have served as scalar rewards for reinforcement learning (RL); we show that they provide substantially richer signal as dense PI for distillation, and contrary to intuition, soft rubric PI provides a larger and more effective training signal on student roll-outs than hard reference completion PI in this regime. A reference completion is one point in a set of valid responses, so distilling towards it over-constrains the student, while rubrics specify the preference structure shared across the set of valid responses. We show the effectiveness of using rubrics as PI for open-ended generation across Qwen and Llama model families and show that it outperforms rubric-as-reward (RaR) RL using HealthBench, a benchmark that grades open-ended health responses against physician-created rubrics, providing dense token-level supervision for open-ended tasks; RuPI beats RaR RL by up to +0.10 absolute score and, under matched recipe and KL direction, beats reference-PI by +0.034 to +0.079 absolute score across three models. We further show that these findings generalize to training on the RubricHub Science corpus and evaluating on ResearchQA: soft rubric PI outperforms both reference-PI distillation and RaR RL (66.6% vs. 64.2% and 57.6%).
发表机构
- Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。