arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习标量奖励模型的端到端潜在推理轨迹

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End

Sanwoo Lee, Clive Bai, Hsiu-Yuan Huang, Kun Liang, Weijie Liu, Yunfang Wu

arXiv 2607.29185首次发表:更新:

AI 中文总结

针对奖励模型范式不匹配问题,提出LatentRM框架,将推理轨迹作为潜在变量端到端优化,在多任务上优于各类基准奖励模型。

AI 中文摘要

奖励模型(RMs)是通过强化学习使大型语言模型与人类偏好对齐的核心组件。传统标量RMs虽能实现高效的概率奖励建模,但依赖表面线索,无法泛化到复杂或分布外(OOD)任务;生成式RMs利用丰富推理提升挑战性任务的鲁棒性,但其基于自然语言的评分缺乏标量RMs具备的数值灵活性与概率可解释性。近期方法虽通过离策略多任务学习结合两种范式,但并行优化无法保证生成的推理轨迹主动对齐或有益于下游标量奖励预测。为解决该不匹配问题,我们提出LatentRM这一奖励建模框架,将中间推理轨迹学习为离散潜在变量,以显式最大化下游标量奖励的似然。通过对潜在推理空间进行端到端的 on-policy 优化,LatentRM 将基于深度推理的评估与精确评分紧密结合。在分布内和OOD数据集及RLHF上的大量验证表明,LatentRM在从开放式对话到复杂推理等各类任务的偏好建模与策略对齐方面,性能优于标量、生成式及混合RMs。

英文摘要

Reward models (RMs) are central to aligning large language models with human preferences via reinforcement learning. Although traditional scalar RMs enable efficient and probabilistic reward modeling, they rely on superficial cues that fail to generalize to complex or out-of-distribution (OOD) tasks. Conversely, generative RMs leverage extensive reasoning to improve robustness on challenging tasks, but their natural language-based scores lack the numerical flexibility and probabilistic interpretability that scalar RMs offer. While recent approaches combine both paradigms through off-policy multi-task learning, such parallel optimization does not guarantee that generated reasoning traces actively align with or benefit downstream scalar reward prediction. To address this mismatch, we propose LatentRM, a reward modeling framework that learns intermediate reasoning traces as discrete latent variables to explicitly maximize the likelihood of downstream scalar rewards. Through on-policy optimization of the latent reasoning space end-to-end, LatentRM tightly couples deep reasoning-based evaluation with precise scoring. Extensive validations on in-distribution and OOD datasets and RLHF show that LatentRM outperforms scalar, generative, and hybrid RMs on preference modeling and policy alignment across tasks ranging from open-ended conversation to complex reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑