arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34695cs.ROcs.AI

零阶奖励塑形对镇定控制中策略梯度的充分性

Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control

  • University of California San Diego(加州大学圣迭戈分校)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Yisheng Zhang, Tao Wang, Sicun Gao

AI总结:

本文研究镇定控制中奖励塑形所需的信息量,证明策略梯度无需一阶(速度)奖励项即可成功,但需零阶完备奖励,为机器人RL奖励设计提供原则性指导。

AI中文摘要:

奖励塑形是现代基于深度强化学习(RL)的机器人控制的基础,然而实践者仍严重依赖从经典最优控制和轨迹优化中借鉴的启发式原则。现有方法很少区分对控制目标内在的奖励项与数值正则化项,导致超参数调优脆弱。为确定奖励必须包含哪些量,我们研究了镇定控制问题,重点关注零阶(构型)和一阶(速度)信息。我们在理论和实验上证明,策略梯度方法可以在没有一阶奖励项的情况下成功解决镇定任务,而添加此类项反而会随着其规模增大而引入严重敏感性。相反,我们的发现证实,奖励函数必须在目标相关坐标上具有零阶完备性,而在我们的低耗散假设下,策略观测中仍需一阶状态。总体而言,这些结果为机器人强化学习中的奖励设计提供了可操作且原则性的指导。

英文摘要:

Reward shaping is fundamental to modern robotic control with deep reinforcement learning (RL), yet practitioners still rely heavily on heuristic principles borrowed from classical optimal control and trajectory optimization. Existing methods rarely distinguish reward terms that are intrinsic to the control objective from numerical regularizers, leading to brittle hyperparameter tuning. To determine which quantities a reward must contain, we study the stabilization control problem with a focus on zeroth-order (configuration) and first-order (velocity) information. We theoretically and empirically demonstrate that policy gradient methods can successfully solve stabilization tasks without first-order reward terms, adding such terms can instead introduce severe sensitivity as their scale grows. Conversely, our findings confirm that reward functions must be zeroth-order complete over goal-relevant coordinates, while the first-order state remains necessary in the policy observation under our low-dissipation assumptions. Overall, these results provide actionable and principled guidance for reward design in robotic RL.

补充信息

↑