发表机构
Fundation Model Research Center, CASIA; School of Artificial Intelligence, UCAS; Institute for AI Industry Research (AIR), Tsinghua University; College of Automotive and Energy Engineering (CAEE), Tongji University(中国科学院自动化研究所基础模型研究中心; 中国科学院大学人工智能学院; 清华大学人工智能产业研究院; 同济大学汽车与能源工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对多奖励RL的对齐税问题,提出PRISM框架,通过策略分解优化正、负策略,在三类任务上优于基线并具备推理可控性。
AI 中文摘要
现代大型语言模型(LLMs)不仅需要正确回答问题,还需使其行为适配不同的人类价值观和用例。因此,多奖励强化学习(RL)已成为LLMs领域愈发重要的问题,其中每个奖励对应期望行为的不同方面。然而,多奖励优化面临更严重的对齐税问题:不同优化目标可能相互权衡甚至冲突,导致后训练不稳定且低效。本研究提出PRISM,一种基于策略空间分解与组合思想的新型多奖励RL框架。PRISM不混合不同奖励,而是优化一组独立的正策略和一个全局负策略,这缓解了多奖励策略优化中的潜在冲突,同时通过灵活的策略组合在推理时实现可控性。在科学推理、工具使用推理及有用性-安全性对齐上的实验表明,PRISM在现有多奖励RL基线中表现始终更优,且具备推理时偏好控制的额外可控性。
英文摘要
Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and use cases. As a result, multi-reward reinforcement learning (RL) has become an increasingly important problem for LLMs, where each reward captures a different aspect of desired behavior. However, optimizing with multiple rewards suffers from a more severe alignment tax issue, where different optimization objectives can trade off or even conflict with each other, leading to unstable and inefficient post-training. In this work, we propose PRISM, a new multi-reward RL framework built upon the idea of policy-space decomposition and composition. Instead of compositing different rewards, PRISM optimizes a set of standalone positive policies and a global negative policy. This alleviates the potential conflict during multi-reward policy optimization, while enabling controllability during inference by flexible policy composition. Experiments on scientific reasoning, tool-use reasoning, and helpfulness-safety alignment show that PRISM consistently outperforms existing multi-reward RL baselines, with extra controllability for inference-time preference control.