arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

奖励更好的思维以实现大语言模型偏好对齐

Rewarding Better Thinking for LLM Preference Alignment

Xubo Liu, Wenya Guo, Ruxue Yan, Xinying Qian, Ying Zhang

arXiv 2607.19824首次发表:更新:

AI 中文总结

研究大语言模型偏好对齐问题,提出思维清单奖励(TCR)方法,将偏好对转化为思维清单评估推理轨迹,引入EMA残差公式减少与结果监督重叠,实验证明该方法能提升对齐性能。

AI 中文摘要

大语言模型偏好对齐旨在针对不同用户指令使模型朝着人类偏好进行优化。强化学习已成为实现此目标的主要训练后方法,但现有代理奖励通常是结果层面的,主要评估最终响应,为推理轨迹提供的指导有限。这会导致在多个响应获得相似最终分数时信用分配粗略,使轨迹层面的偏好未得到充分明确。为解决这一限制,我们提出思维清单奖励(TCR),一种用于基于强化学习的偏好对齐的面向过程的奖励。TCR将偏好对转换为特定样本的思维清单,并用它们评估生成的推理轨迹是否解决了偏好隐含的考虑因素。为减少与结果层面监督的重叠,TCR进一步引入指数移动平均(EMA)残差公式,以分离出超出结果奖励可预测范围的互补思维盈余。对来自三个模型家族的五个模型进行的实验表明,TCR在不同基准上持续提高对齐性能,消融实验进一步验证了基于EMA的残差公式和特定样本清单监督的重要性。

英文摘要

LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has become a major post-training approach for this goal, but existing proxy rewards are often outcome-level, mainly evaluating the final response while providing limited guidance for the reasoning trajectory. This can make credit assignment coarse when multiple responses receive similar final scores, leaving trajectory-level preferences under-specified. To address this limitation, we propose Thinking Checklist Reward (TCR), a process-oriented reward for RL-based preference alignment. TCR converts preference pairs into sample-specific thinking checklists and uses them to evaluate whether the generated reasoning trace addresses the preference-implied considerations. To reduce overlap with outcome-level supervision, TCR further introduces an exponential moving average (EMA) residual formulation to isolate a complementary thinking surplus beyond what is predictable from the outcome reward. Experiments on five models from three model families show that TCR consistently improves alignment performance across diverse benchmarks, with ablations further validating the importance of EMA-based residual formulation and sample-specific checklist supervision.

Commentsunder review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑