从语言反馈中通过位置选择性自蒸馏训练LLM裁判
Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
浏览论文内容
中文总结 AI 辅助
本研究提出位置选择性自蒸馏方法,通过熵移掩蔽利用语言反馈训练LLM裁判,在主观任务上比结果监督RL提升2-9个百分点,并改善分布外泛化。
中文摘要 AI 辅助
我们研究从自然语言反馈中训练LLM裁判,特别是针对主观任务,其中判决在很大程度上取决于裁判引用哪些评估标准以及如何权衡这些标准。主导方法,即结果监督强化学习(例如GRPO),为rollout中的每个token赋予一个仅由最终判决准确性决定的单一标量,没有在标准选择token处提供单独的信用,并且忽略了伴随偏好标签自然出现的丰富语言反馈(例如偏好理由)。自蒸馏(SD)是利用这种语言反馈的一种自然方式:同一模型,以该反馈为条件,充当提供密集、位置级监督的教师。然而,并非所有位置都携带同等有用的信号。利用教师和学生之间的每位置熵移,我们识别出两种机制:上下文锐化,其中教师将概率集中在特定的反馈对齐标准表达上;以及上下文扩散,其中教师将概率分布在多个反馈对齐的备选方案上。我们将这些模式解释如下:锐化鼓励记忆特定的标准表达,而扩散通过保留这些备选方案来促进语义理解。受这种不对称性的启发,我们引入了基于熵移的位置掩蔽,保留熵移分布的下尾。实验表明,掩蔽较高熵移位置比朴素SD改善了分布外泛化。由此产生的自蒸馏裁判在评估的主观子类别上比结果监督RL训练的裁判高出2-9个百分点,同时在客观子类别上保持竞争力。
英文摘要
We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) that naturally accompanies preference labels. Self-Distillation (SD) is one natural way to use this language feedback: the same model, conditioned on this feedback, acts as a teacher providing dense, position-level supervision. However, not all positions carry equally useful signal. Using the per-position entropy shift between teacher and student, we identify two regimes: context sharpening, where the teacher concentrates probability on a particular feedback-aligned criterion expression, and context spreading, where the teacher distributes probability across multiple feedback-aligned alternatives. We interpret these patterns as follows: sharpening encourages memorization of a particular criterion expression, whereas spreading promotes semantic understanding by preserving these alternatives. Motivated by this asymmetry, we introduce position masking based on the entropy shift that retains the lower tail of the entropy-shift distribution. Experiments show that masking higher-entropy-shift positions improves out-of-distribution generalization over naive SD. The resulting self-distilled judges outperform judges trained with outcome-supervised RL by 2-9 percentage points on the evaluated subjective subcategories, while remaining competitive on objective ones.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
- Amazon(亚马逊)
- UC San Diego(加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。