发表机构
Lancaster University; Beihang University; Tsinghua University(兰卡斯特大学; 北京航空航天大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现熵的测量位置会影响PPO学习的有界连续控制策略几何,通过MyoLeg、Dog-Stand任务实验,表明熵测量空间是均值-方差耦合的设计选择,仅任务回报无法表征有界策略几何。
AI 中文摘要
许多连续控制策略被优化为无界高斯分布后再映射到有界动作空间。我们证明,熵的测量位置会改变近端策略优化(PPO)学习到的策略几何。在80肌肉的MyoLeg任务中,裁剪高斯策略执行的动作中有89.07%落在边界的5%范围内。相同状态分解显示,这并非仅由方差导致:将方差设为零后,仍有83.83%的动作靠近边界,同时82.12%的状态条件均值位于可执行区间之外。用tanh映射替代裁剪无法消除高方差状态。对于潜在高斯熵H(u),熵损失对均值的梯度为零,对方差的梯度为常数;对于执行动作的熵H(a),变换雅可比会在均值上添加向内的梯度。在三个匹配的MyoLeg随机种子下,潜在熵、无熵、执行动作熵对应的近边界占比分别为71.42%、29.76%和18.83%。采用独立CleanRL-based PPO实现的38维Dog-Stand复现了均值几何的排序,该排序在共享状态评估和1%至10%的边界裕度下仍成立。直接均值惩罚可匹配或超过H(a)产生的中心化效果,表明内部均值并非执行熵独有。然而,匹配的均值几何可伴随显著不同的方差和回报。因此,熵测量空间是均值-方差耦合的设计选择,仅任务回报无法表征有界策略几何。
英文摘要
Many continuous-control policies are optimized as unbounded Gaussians and then mapped into bounded actions. We show that where entropy is measured changes the policy geometry learned by proximal policy optimization (PPO). In an 80-muscle MyoLeg task, a clipped Gaussian executes 89.07% of actions within 5% of a bound. A same-state decomposition shows that this is not due to variance alone: setting variance to zero still leaves 83.83% of actions near a bound, while 82.12% of state-conditioned means lie outside the executable interval. Replacing clipping with a tanh map does not remove the high-variance regime. For latent Gaussian entropy H(u), the entropy loss has zero gradient with respect to the mean and a constant variance-increasing gradient. For executed-action entropy H(a), the transform Jacobian adds an inward gradient on the mean. Across three matched MyoLeg seeds, near-boundary occupancy is 71.42%, 29.76%, and 18.83% under latent entropy, no entropy, and executed-action entropy. A 38-dimensional Dog-Stand replication with an independent CleanRL-based PPO implementation reproduces the ordering in mean geometry, which also survives shared-state evaluation and boundary margins from 1% to 10%. Direct mean penalties can match or exceed the centering produced by H(a), showing that interior means are not unique to executed entropy. However, matched mean geometry can coexist with substantially different variance and return. Entropy measurement space is therefore a coupled mean-variance design choice, and task return alone does not characterize bounded-policy geometry.
Comments24 pages, 6 figures, 8 tables