arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于约束控制的结合排名增强与自适应优势归一化的同策略优化

Ranking-Augmented On-Policy Optimization with Adaptive Advantage-Normalization for Constrained Control

Md Ragib Rownak, Sidra Ghayour Bhatti, Qadeer Ahmed

arXiv 2608.15359首次发表:更新:

AI 中文总结

本文提出A-GRPO方法,通过轨迹级排名机制与自适应优势归一化优化约束控制,在3605步混合动力能量管理任务中,其持续可行性达75.4%,优于PPO-Lag基线。

AI 中文摘要

本文分析了Advantage-Ranked Group Relative Policy Optimization(A-GRPO,一种结合排名增强、无评论家的策略梯度方法,采用Transformer编码器作为动作器,用于带终端约束的固定时域控制)的有界性和可行性特性。当仅在最终步骤评估可行性时,产生的稀疏反馈会破坏基于评论家的优势估计,并削弱标准拉格朗日方法。本文提出了一种轨迹级排名机制,该机制通过根据约束满足情况对优势进行重新加权,来增强组相对策略更新,并确立了三个结果:(i)一种尺度自适应的每时间步归一化方法,可独立约束每个时间步的优势方差;(ii)在可验证的排名权重条件下,排名后的优势能严格区分可行轨迹与违反约束的轨迹,使策略梯度偏向约束满足;(iii)自适应对偶变量保持有界,且呈现漂移平衡特性,可作为可行性的反馈机制。这些结果在一个3605步的串联混合动力动力总成能量管理任务(带有终端荷电状态约束)上得到验证,其中A-GRPO实现了75.4%的平均持续可行性,回报与动态规划最优值的差距在3.7%以内,优于带拉格朗日惩罚的近端策略优化(PPO-Lag)基线(27.4%的持续可行性),消融实验证实排名组件和拉格朗日组件对该性能均是必要的。

英文摘要

This paper analyzes the boundedness and feasibility properties of Advantage-Ranked Group Relative Policy Optimization (A-GRPO), a ranking-augmented, critic-free policy gradient method employing a Transformer-encoder actor for fixed-horizon control with terminal constraints. When feasibility is evaluated only at the final step, the resulting sparse feedback destabilizes critic-based advantage estimation and weakens standard Lagrangian approaches. A trajectory-level ranking mechanism that augments group-relative policy updates by reweighting advantages according to constraint satisfaction is formalized, and three results are established: (i) a scale-adaptive per-timestep normalization bounds advantage variance at every timestep independently, (ii) the ranked advantage strictly separates feasible from violating trajectories under a verifiable ranking-weight condition, biasing the policy gradient toward constraint satisfaction, and (iii) the adaptive dual variables remain bounded and exhibit a drift-balance property that acts as a feedback mechanism for feasibility. These results are validated on a 3,605-step series-hybrid powertrain energy management task with a terminal state-of-charge constraint, where A-GRPO achieves 75.4% mean sustained feasibility with return within 3.7% of the dynamic programming optimum, outperforming a Proximal Policy Optimization with Lagrangian penalties (PPO-Lag) baseline (27.4% sustained), and ablation experiments confirm that both the ranking and Lagrangian components are necessary for this performance.

Comments8 pages, 2 figures. Accepted for presentation at the 2026 IEEE Conference on Decision and Control (CDC). (c) 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media. The full copyright notice appears on the first page of the paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑