稳定价值冲突解决的几何视角
A Geometric Perspective on Stabilizing Value Conflict Resolution
- University of Illinois - Urbana-Champaign(伊利诺伊大学厄巴纳 - 香槟分校)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究大语言模型在RLHF训练中应对价值冲突的问题,核心方法是利用思维链推理,通过几何视角分析其作用,创建新的以价值冲突为重点的CoT设计,提升了道德推理性能,为改进模型处理复杂价值冲突请求的性能提供新途径。
AI中文摘要:
大语言模型(LLMs)在通过人类反馈强化学习(RLHF)的压缩标量奖励进行训练时,常常难以应对价值冲突。为应对这一挑战,我们研究了思维链(CoT)推理如何有助于提升该领域的性能。从几何角度看,我们表明CoT与在模型损失景观最陡峭方向上进一步平滑相关,有助于解决传统标量奖励的优化不稳定性。我们还通过相关下游基准证明,以价值冲突为重点的CoT可能推广到不同类型的道德推理,表明这种CoT有潜力成为更好道德推理的有效机制。为利用这一潜力,我们创建了一种新的以价值冲突为重点的CoT设计,进一步平滑损失景观的最陡峭方向并提高道德推理性能。这一发现表明,明确修改和改进推理动态设计为提升模型在复杂价值冲突用户请求上的性能提供了一条有前景的途径,推动了大语言模型的多元对齐。
英文摘要:
Large Language Models (LLMs) often struggle to navigate value conflicts when trained with the compressed scalar rewards of Reinforcement Learning from Human Feedback (RLHF). To address this challenge, we investigate how chain-of-thought (CoT) reasoning can help improve performance in this domain. Geometrically, we show that CoT correlates with further smoothing the model's loss landscape in its sharpest direction, helping resolve the optimization instability of traditional scalar rewards. We also demonstrate via relevant downstream benchmarks that value conflict-focused CoT may generalize to different kinds of moral reasoning, demonstrating that this CoT has the potential to be an effective mechanism for better moral reasoning. To capitalize on this potential, we create a new value conflict-focused CoT design that further smooths the sharpest direction of the loss landscape and increases moral reasoning performance. This finding shows that explicitly modifying and improving the design of reasoning dynamics offers a promising avenue for improving model performance on user requests with complex value conflicts, advancing pluralistic alignment in LLMs.