arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

凸-凹强化学习

Convex-Concave Reinforcement Learning

Shripad V. Deshmukh, Yaswanth Chittepu, Dhawal Gupta, Philip Thomas, Scott Niekum

arXiv 2610.09108首次发表:更新:

AI 中文总结

本文提出凸-凹强化学习(CCRL),将策略优化转化为DC约束DC规划,统一CPI、NPG、TRPO、AWR,并通过多步轴提升信用传播,在诊断MDP和医疗保健任务中优于PPO。

AI 中文摘要

策略学习驱动着当今强化学习中许多最具影响力且投资巨大的应用。然而,其所依赖的核心优化问题(最大化期望回报)以非凸著称,即使在直接策略参数化下也是如此,该领域主要通过回避这一问题来应对:在信任域约束下优化回报的凸代理近似(NPG、TRPO、PPO、AWR)。我们表明,这个看似无结构的问题实际上并非无结构。在对数密度比坐标 $y:= \log[\pi/\pi_n]$ 中,精确的每轮迭代目标(可通过逐决策重要性采样(PDIS)计算)是一个受凸-凹约束的凸-凹差分(DC-constrained DC)规划。这一结构使我们能够超越代理近似:它沿可解释的轴将CPI、NPG、TRPO和AWR恢复为特例,并开辟了一个耦合连续决策的多步轴 $k$。我们使用序贯凸规划(SCP)——求解凸-凹差分问题的标准求解器——来解决每轮迭代规划,并在温和条件下给出收敛保证,从而弥合了凸-凹差分优化与强化学习文献之间的鸿沟。实验上,多步凸-凹强化学习(CCRL)在信用必须跨时间范围传播的诊断性MDP上获胜(其优势随依赖长度增长),在经典控制上与调优的PPO相当,并且在一个现实的、随机的、中期时间范围的医疗保健领域中,收敛到相同的近最优生存率的速度明显快于调优的PPO,训练曲线下面积高出11.3%。

英文摘要

Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even under a direct policy parameterization, and the field has largely responded by avoiding it: optimizing convex surrogate approximations of the return under trust-region constraints (NPG, TRPO, PPO, AWR). We show that this seemingly unstructured problem is not actually structureless. In log-density-ratio coordinates $y := \log[π/π_n]$, the exact per-iteration objective, computable via per-decision importance sampling (PDIS), is a difference-of-convex-constrained difference-of-convex (DC-constrained DC) program. This structure lets us move beyond surrogate approximations: it recovers CPI, NPG, TRPO, and AWR as special cases along interpretable axes, and it opens a multi-step axis $k$ that couples consecutive decisions. We solve the per-iteration program with sequential convex programming (SCP), the standard solver for difference-of-convex problems, and give convergence guarantees under mild conditions, bridging the difference-of-convex optimization and RL literatures. Empirically, multi-step Convex-Concave RL (CCRL) wins on diagnostic MDPs where credit must propagate across a horizon (its advantage growing with the dependency length), is competitive with a tuned PPO on classic control, and on a realistic, stochastic, mid-horizon healthcare domain converges markedly faster than tuned PPO to the same near-optimal survival, with an 11.3% higher area under the training curve.

Comments37 pages, 5 figures. Code: https://github.com/UMass-SCALAR-Lab/Convex-Concave-RL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑