arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不要混合奖励,要混合策略:多奖励强化学习的策略分解与优化

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

Ruiming Liang, Yi Zhong, Yizhen Yuan, Yinan Zheng, Tianyi Tan, Tianyue Wang, Haiyun Guo, Jinqiao Wang, Xianyuan Zhan

arXiv 2607.29246首次发表:更新:

发表机构

Fundation Model Research Center, CASIA; School of Artificial Intelligence, UCAS; Institute for AI Industry Research (AIR), Tsinghua University; College of Automotive and Energy Engineering (CAEE), Tongji University(中国科学院自动化研究所基础模型研究中心; 中国科学院大学人工智能学院; 清华大学人工智能产业研究院; 同济大学汽车与能源工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对多奖励RL的对齐税问题,提出PRISM框架,通过策略分解优化正、负策略,在三类任务上优于基线并具备推理可控性。

AI 中文摘要

现代大型语言模型(LLMs)不仅需要正确回答问题,还需使其行为适配不同的人类价值观和用例。因此,多奖励强化学习(RL)已成为LLMs领域愈发重要的问题,其中每个奖励对应期望行为的不同方面。然而,多奖励优化面临更严重的对齐税问题:不同优化目标可能相互权衡甚至冲突,导致后训练不稳定且低效。本研究提出PRISM,一种基于策略空间分解与组合思想的新型多奖励RL框架。PRISM不混合不同奖励,而是优化一组独立的正策略和一个全局负策略,这缓解了多奖励策略优化中的潜在冲突,同时通过灵活的策略组合在推理时实现可控性。在科学推理、工具使用推理及有用性-安全性对齐上的实验表明,PRISM在现有多奖励RL基线中表现始终更优,且具备推理时偏好控制的额外可控性。

英文摘要

Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a result, multi-reward reinforcement learning (RL) has become an increasingly important problem for LLMs, where each reward captures a different aspect of desired behavior. However, optimizing with multiple rewards suffers from a more severe alignment tax issue, where different optimization objectives can trade off or even conflict with each other, leading to unstable and inefficient post-training. In this work, we propose PRISM, a new multi-reward RL framework built upon the idea of policy-space decomposition and composition. Instead of compositing different rewards, PRISM optimizes a set of standalone positive policies and a global negative policy. This alleviates the potential conflict during multi-reward policy optimization, while enabling controllability during inference by flexible policy composition. Experiments on scientific reasoning, tool-use reasoning, and helpfulness-safety alignment show that PRISM consistently outperforms existing multi-reward RL baselines, with extra controllability for inference-time preference control.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑