arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27572cs.LGcs.AI

DCRL:通过策略-奖励流形对齐的解耦与耦合强化学习

DCRL: Decoupling and Coupling Reinforcement Learning via Policy-Reward Manifold Alignment

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Tencent HY(腾讯HY)

机构由 AI 辅助整理,请以论文原文为准。

Henan Sun, Zehua Li, Haitao Hu, Qifan Zhang, Jianfeng Zhang, Nuo Chen, Jia Li

AI总结:

提出DCRL框架,通过策略-奖励流形对齐解决RL中奖励系统不稳定与奖励黑客问题,利用逻辑提示演化和联合更新机制,显著提升LLM推理性能,小模型超越大模型。

AI中文摘要:

强化学习(RL)已成为提升大型语言模型(LLMs)推理能力的关键范式。然而,现有的奖励系统,如基于规则的和基于奖励模型的系统,常常表现出优化不稳定和奖励黑客等问题。在这项工作中,我们从几何视角重新审视LLMs的一般推理,将其概念化为一个由三个相互依赖的子流形组成的耦合流形:逻辑演绎、评估和表示。基于这一视角,RL中的响应生成可以被解释为从评估流形解耦的过程,而奖励估计则对应于从逻辑演绎流形解耦的过程。基于规则和基于奖励模型的RL系统的局限性可以在几何上解释为RL过程中策略-奖励流形的不匹配。为了解决上述错位问题,我们提出了解耦与耦合强化学习(DCRL)框架,该框架包含两个关键组件:(1)基于三段论逻辑的提示演化机制,动态细化奖励标准以增强奖励流形的表达能力;(2)策略-奖励重新耦合机制,联合更新奖励和策略模型,确保评估一致性并缓解训练过程中的流形不匹配。理论分析和跨多个推理领域的广泛实验表明,DCRL始终优于基于规则和基于奖励模型的基线。值得注意的是,在DCRL下训练的Qwen3-4B模型超越了Qwen3-32B基线,并接近Qwen3-235B模型的性能,突显了其在RL中的卓越有效性和泛化能力。

英文摘要:

Reinforcement learning (RL) has emerged as a key paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing reward systems, such as rule-based and reward-model-based, often exhibit issues such as unstable optimization and reward hacking. In this work, we revisit the general reasoning of LLMs from a geometric perspective, conceptualizing it as a coupled manifold composed of three interdependent sub-manifolds: logical deduction, evaluation, and representation. Based on this perspective, response generation in RL can be interpreted as a decoupling process from the evaluation manifold, while reward estimation corresponds to a decoupling process from the logical deduction manifold. The limitations of rule-based and reward-model RL systems can be geometrically interpreted as the mismatch of policy-reward manifolds during RL process. To address the aforementioned misalignment, we propose Decoupling and Coupling Reinforcement Learning (DCRL) framework, which incorporates two key components: (1) a syllogistic logic-based prompt evolution mechanism that dynamically refines reward rubrics to enhance the expressiveness of the reward manifold; and (2) a policy-reward re-coupling mechanism that jointly updates the reward and policy models, ensuring consistent evaluation and mitigating manifold mismatch during training. Theoretical analysis and extensive experiments across multiple reasoning domains demonstrate that DCRL consistently outperforms both rule-based and reward-model baselines. Notably, a Qwen3-4B model trained under DCRL surpasses a Qwen3-32B baseline and approaches the performance of a Qwen3-235B model, highlighting superior effectiveness and generalization in RL.

补充信息

↑