arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

P3:基于稳定变分自编码器的机器人学习的概率策略传播

P3: Probabilistic Policy Propagation for Stable VAE-Based Robot Learning

Liyun Yan, Jianming Ma, Yang Zhang, Shengcheng Fu, Zhanxiang Cao, Keqi Zhu, Yizhi Chen, Yue Gao

arXiv 2607.25541首次发表:更新:

AI 中文总结

研究基于变分自编码器的机器人学习中随机潜变量与近端策略优化不匹配问题,提出P³框架,它结合概率方法与采样校准,实验表明可提高数据效率、减少收敛步数,为人形跑酷任务的基于变分自编码器的PPO奠定基础。

AI 中文摘要

变分自编码器在机器人技术中广泛用于对高维且有噪声的观测进行编码。然而,其随机潜变量与近端策略优化(PPO)不匹配:有效策略对潜变量分布求边缘概率,而之前的实现仅用一个潜变量样本估计概率比和KL散度。我们发现一个根本但被忽视的理论原因:随机潜变量空间中的朴素单样本近似在替代损失中会导致显著的方差和偏差。为解决此问题,我们引入P³(概率策略传播),这是一个基于变分自编码器策略的分布感知优化框架。P³将基于矩的概率方法用于稳定高效学习,并用基于采样的校准来应对潜变量不确定性下的稳健策略行为。在实验中,P³将数据效率从64.6%提高到>96%,收敛步数减少>20%。此外,P³在具有挑战性的人形跑酷任务上进行了评估,为基于变分自编码器的PPO奠定了有效基础。代码可从此https URL获取。

英文摘要

Variational Autoencoders are widely used to encode high-dimensional and noisy observations in robotics. However, their stochastic latent creates a mismatch with Proximal Policy Optimization (PPO): an effective policy marginalizes over the latent distribution, whereas former implementations estimate its probability ratio and KL divergence using only one latent sample. We identify a fundamental but overlooked theoretical cause: naive single-sample approximations in stochastic latent space induce significant variance and bias in the surrogate loss. To address this, we introduce P^3 (Probabilistic Policy Propagation), a distribution-aware optimization framework for VAE-based policies. $P^3$ couples moment-based probabilistic method for stable and efficient learning with sampling-based calibration for robust policy behavior under latent uncertainty. In our experiments, P^3 boosts data efficiency from 64.6% to >96%, reduces convergence steps by >20%. Furthermore, P^3 is evaluated on challenging humanoid parkour tasks and shows an effective foundation for VAE-based PPO. Code is available at https://github.com/ylyem9x/P3_Open.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑