arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

S2T-RLHF:基于稳定偏好的强化学习从人类反馈中的分层信用分配

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF

Wei Chen, Guanghui Zhu, Yafei Li, Limin Wang, Yihua Huang

arXiv 2607.18258首次发表:更新:

发表机构

State Key Laboratory for Novel Software Technology, Nanjing University; School of Computer Science and Artificial Intelligence, Zhengzhou University(南京大学计算机软件新技术国家重点实验室; 郑州大学计算机科学与人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于偏好奖励模型的RLHF训练不稳定问题,提出分层信用分配的粒度感知原则,引入S2T-RLHF框架,先在句子间分配奖励,再在句内细化,无需重新训练模型,实验证明该方法提高了训练稳定性和鲁棒性。

AI 中文摘要

基于偏好奖励模型的人类反馈强化学习(RLHF)通常呈现不稳定的训练动态。一个关键因素是标准RLHF依赖单一序列级标量奖励,传播到令牌级策略更新时,响应内的信用分配本质上模糊不清。近期工作试图通过将奖励细化为更密集的令牌级监督来解决此问题,但该假设不完整。我们提出分层信用分配的粒度感知原则,强调面向稳定性的奖励设计。基于此,引入S2T-RLHF,它先在句子间分配序列级偏好奖励,再在句内进行有界令牌级细化,无需重新训练奖励模型或令牌级监督。实验表明,S2T-RLHF提高了训练稳定性和鲁棒性,同时保持了有竞争力的偏好对齐。

英文摘要

Reinforcement learning from human feedback (RLHF) with preference-based reward models often exhibits unstable training dynamics. A key contributing factor is that standard RLHF relies on a single sequence-level scalar reward, which is propagated to token-level policy updates and leaves credit assignment within a response inherently ambiguous. Recent work has attempted to address this issue by refining rewards into denser token-level supervision, often relying on the implicit assumption that finer-grained credit assignment improves optimization. We argue that this assumption is incomplete: when preference signals are noisy and only defined at the response level, overly fine-grained reward refinement can amplify reward uncertainty and destabilize learning. To address this problem, we propose a granularity-aware principle for hierarchical credit assignment, emphasizing stability-oriented reward design rather than maximal allocation precision. Under this principle, sentences serve as a natural intermediate granularity, balancing semantic coherence with robustness to token-level noise. Guided by this view, we introduce S2T-RLHF. This sentence-to-token reward decomposition framework first allocates sequence-level preference rewards across sentences and then applies bounded token-level refinement within each sentence, without reward-model retraining or token-level supervision. Experiments across multiple datasets and optimization settings show that S2T-RLHF improves training stability and robustness while maintaining competitive preference alignment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑