arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向强化学习的鲁棒通用效用方法

Robust General Utility for Reinforcement Learning

Zixuan Liu, Fangzheng Wu, Brian Summa, Zizhan Zheng

arXiv 2608.03562首次发表:更新:

AI 中文总结

针对通用效用强化学习的部署效用误指定问题,提出极小极大学习框架及两种收敛随机算法,经LLM安全对齐等实验验证了其有效性。

AI 中文摘要

带通用效用的强化学习(RL)通过优化策略诱导占用度测度的任意效用函数,扩展了经典RL,从而支持更广泛的应用。然而,此前的通用效用RL研究通常假设评估效用是固定且正确指定的。在实际部署中,使用的效用可能与训练时的效用存在偏差,产生了现有研究未解决的鲁棒性差距。基于此,我们提出鲁棒通用效用RL,这是一种极小极大学习框架,可在规定的不确定性集内针对效用误指定训练策略。我们的框架严格泛化了标准通用效用RL,还通过效用不确定性集的适当选择,为许多现有RL框架(包括奖励鲁棒RL和约束RL)提供了统一视角。我们进一步针对两种场景开发了可证收敛的随机算法:对于凹效用,我们提出投影随机梯度下降上升法,并建立了平稳性保证;对于更具挑战性的非凹场景,我们提出随机近邻额外梯度算法,该算法可缓解非凹性导致的不适定行为,且具有近似一阶平稳性的收敛保证。在LLM安全对齐和探索最大化任务上的实验进一步证实了与我们理论一致的收敛行为。

英文摘要

Reinforcement learning (RL) with general utility extends classic RL by optimizing an arbitrary utility functional of the policy-induced occupancy measure, thereby enabling a broader range of applications. However, previous work on general utility RL typically assumes the evaluation utility is fixed and correctly specified. In practice, the utility used at deployment can deviate from the training one, creating a robustness gap that prior work does not address. Motivated by this, we propose robust general-utility RL, a minimax learning framework that trains policies against utility misspecification within a prescribed uncertainty set. Our framework strictly generalizes standard general-utility RL while also providing a unified view of many existing RL frameworks, including reward-robust RL and constrained RL, through appropriate choices of the utility uncertainty set. We further develop provably convergent stochastic algorithms for two regimes. For concave utilities, we develop a projected stochastic gradient descent-ascent method and establish stationarity guarantees. For the more challenging nonconcave regime, we propose a stochastic prox-extragradient algorithm that mitigates ill-posed behavior induced by nonconcavity, with convergence guarantees to approximate first-order stationarity. Experiments on LLM safety alignment and exploration maximization tasks further corroborate the convergence behavior consistent with our theory.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑