arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提升的贝尔曼线性规划用于离线强化学习

Lifted Bellman Linear Programming for Offline Reinforcement Learning

Hyukjun Yang, Jongchan Park, Narim Jeong, Donghwan Lee

arXiv 2609.24489首次发表:更新:

发表机构

NAVER LABS(NAVER 实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出提升的贝尔曼线性规划(LBLP)及其近似算法ALBUM,通过不等式约束实现样本内贝尔曼最优性,无需目标网络,在OGBench上以最少参数和内存达到与FQL相当的性能。

AI 中文摘要

离线强化学习(RL)通常通过最小化回归损失来训练评论员(critic),该损失针对由指数移动平均(EMA)更新的目标网络稳定的自举价值目标。多步目标包含行为策略的动作,因此需要离策略校正。我们转而通过不等式约束在评论员上施加样本内贝尔曼最优性。我们提出了提升的贝尔曼线性规划(LBLP),它将贝尔曼最优性的线性规划表征提升到联合$(Q,V)$空间,使得每个约束仅涉及数据集中的状态-动作对。其唯一最小化器是样本内最优对,并且沿数据集轨迹的$K$步段的约束对于任何 rollout 策略和视界都保持该最小化器不变。在确定性动力学下,该最小化器位于最佳数据集回报与最优值之间。将约束松弛为铰链惩罚,在表格情形下,当惩罚系数有限时恢复相同的解。近似提升的贝尔曼无约束最小化(ALBUM)使用神经网络实现此松弛,并通过停止梯度分离$K$步 rollout 目标。其目标函数不包含对自举目标的平方回归,因此可以在没有目标网络或EMA更新的情况下训练。在确定性动力学下,LBLP解是分离更新的驻点,其系数条件独立于$\gamma$和$K$,并且不等式约束允许沿数据集轨迹的折扣回报作为下界,无需离策略校正或动作分块。在OGBench上,ALBUM使用带有高斯策略的单个评论员,匹配FQL的平均性能,并与最近的动作分块方法相当,同时在所有比较方法中使用最少的参数和最低的峰值GPU内存。

英文摘要

Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions and therefore require off-policy correction. We instead impose in-sample Bellman optimality on the critic through inequality constraints. We formulate the Lifted Bellman Linear Program (LBLP), which lifts the linear programming characterization of Bellman optimality to the joint $(Q,V)$ space so that every constraint involves only state-action pairs in the dataset. Its unique minimizer is the in-sample optimal pair, and constraints along $K$-step segments of dataset trajectories leave this minimizer unchanged for any rollout policy and horizon. Under deterministic dynamics, this minimizer lies between the best dataset return and the optimal value. Relaxing the constraints into hinge penalties recovers the same solution above a finite penalty coefficient in the tabular case. Approximate Lifted Bellman Unconstrained Minimization (ALBUM) implements this relaxation with neural networks and detaches the $K$-step rollout targets by stop gradient. Its objective contains no squared regression onto bootstrapped targets, so it can be trained without target networks or EMA updates. Under deterministic dynamics, the LBLP solution is a stationary point of the detached update under a coefficient condition independent of $γ$ and $K$, and the inequality constraints allow discounted returns along dataset trajectories to serve as lower bounds without off-policy correction or action chunking. On OGBench, ALBUM uses a single critic with a Gaussian policy, matches the average performance of FQL, and is comparable to recent action-chunking methods, while using the fewest parameters and the least peak GPU memory among all compared methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑