发表机构
Harbin Institute of Technology; JD.com Inc.; Institute of Automation, Chinese Academy of Sciences(哈尔滨工业大学; 京东集团; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多目标强化学习中的稀疏奖励与冲突问题,提出多边际偏好优化(MMPO),在数据、梯度和约束层面干预,提升稳定性并泛化至工具学习与代码生成。
AI 中文摘要
现实世界中的多目标强化学习(MORL)常常面临稀疏奖励、奖励冲突以及后期奖励拉锯战等问题,导致传统的线性标量化方法出现严重的指标振荡。为了解决现实部署场景中多个目标之间的优化冲突,我们提出了多边际偏好优化(MMPO),这是一种细粒度的框架,在数据、梯度和约束层面进行干预,而非依赖粗粒度的全局标量化。具体而言,MMPO 通过暴露去偏来缓解稀疏且有偏的奖励,应用优先级感知的正交投影来解耦冲突的梯度,并引入自提示梯度约束以防止主导目标压过较弱目标。在真实电子商务数据集上的实验表明,MMPO 提升了训练稳定性,并在冲突指标上持续取得更优性能。此外,它还能稳健地泛化到更广泛的任务,如 ToolRL 和代码生成,证明了其作为多目标对齐实用范式的有效性。
英文摘要
Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward conflicts, and late-stage reward tug-of-war, causing traditional linear scalarization to experience severe metric oscillations. To address optimization conflicts among multiple objectives in real-world deployment scenarios, we propose Multi-Marginal Preference Optimization (MMPO), a fine-grained framework that intervenes at the data, gradient, and constraint levels rather than relying on coarse-grained global scalarization. Specifically, MMPO performs exposure debiasing to mitigate sparse and biased rewards, applies priority-aware orthogonal projection to decouple conflicting gradients, and introduces self-prompted gradient constraints to prevent dominant objectives from overwhelming weaker ones. Experiments on real-world e-commerce datasets show that MMPO improves training stability and consistently achieves better performance across conflicting metrics. Moreover, it generalizes robustly to broader tasks such as ToolRL and code generation, demonstrating its effectiveness as a practical paradigm for multi-objective alignment.
CommentsAccepted to EMNLP 2026