注意相位:腿式运动中的有效秩与表示健康
Mind the Phase: Effective Rank and Representation Health in Legged Locomotion
浏览论文内容
中文总结 AI 辅助
本研究通过策略雅可比矩阵的有效秩,按步态相位分析腿式运动策略的表示健康,发现层归一化和残差连接为摆动相分配更多秩维度,并提出简单方法实现约3倍更低的关节抖动,改善模拟到现实的迁移。
中文摘要 AI 辅助
强化学习已成为腿式运动中的主导范式,通过大规模并行模拟实现了从后空翻到跑酷等复杂行为。在PPO的非平稳性下,浅层网络仍是事实上的架构,并辅以精心分阶段的课程和环境,然而这些策略学习到的表示仍缺乏深入理解,导致在训练时没有信号可预测其在硬件上的表现。在本工作中,我们通过策略雅可比矩阵的有效秩对运动策略进行实证研究,并表明将秩按步态相位进行条件化能够揭示全局秩平均所掩盖的架构结构。具体而言,我们发现标准架构选择,即层归一化和残差连接,将有效秩的约两个维度分配给摆动相而非支撑相,而这一现象在普通MLP中完全不存在。基于此,我们提出一个简单方法,将这些表示特征转化为更平滑、更可靠的模拟到现实迁移。在实践中,这导致关节抖动降低约3倍,且从仿真到物理Spot机器人都保持一致,表明表示健康是跟踪模拟到现实平滑度的有效训练时视角。
英文摘要
Reinforcement learning has become the leading paradigm in legged locomotion, enabling complex behaviors from backflips to parkour through massively parallel simulation. Under PPO's non-stationarity, shallow networks remain the de facto architecture, supported by carefully staged curricula and environments, yet the representations these policies learn stay poorly understood, leaving no training-time signal of how they will behave on hardware. In this work, we empirically study locomotion policies through the effective rank of the policy Jacobian and show that conditioning rank on the gait phase exposes architectural structure that global rank averages away. In particular, we find that standard architectural choices, namely layer normalization and residual connections, allocate roughly two more dimensions of effective rank to swing than to stance, which is fully absent in vanilla MLPs. Building on this, we propose a simple recipe that turns these representational signatures into smoother, more reliable sim-to-real transfer. In practice, this results in roughly 3x lower joint jitter that holds from simulation onto a physical Spot, suggesting that representation health is an effective training-time lens to track sim-to-real smoothness.
发表机构
- University of São Paulo(圣保罗大学)
机构由 AI 辅助整理,请以论文原文为准。