arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SleepWalking:用于四足机器人端到端盲运动的特权表征塑造方法

SleepWalking: Privileged Representation Shaping for End-to-End Blind Locomotion in Legged Robots

Zheng Pan, Tenghui Wang, Peilin Li, Shiyu Zhou, Hao Sun, Yan Ma, Liang Yu, Liang He

arXiv 2608.30883首次发表:更新:

发表机构

Northwestern Polytechnical University; Shanghai Jiao Tong University; Yunmu Intelligent Manufacturing Co., Ltd.(西北工业大学; 上海交通大学; 云沐智能制造有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出SWAQ框架,通过特权物理重构塑造历史表征,在四足机器人盲运动任务中,相比DWAQ提升地形水平同时降低推理计算量,为部分可观测运动学习提供新方法。

AI 中文摘要

部分可观测的运动任务要求策略在机器人与环境状态中与任务相关的属性未被瞬时观测完全指定时仍能执行。现有方法通常通过显式估计缺失的物理变量,或通过结构化架构处理扩展观测历史来应对这一挑战。我们提出不同观点:部分可观测性本质上是信息保留问题,关键问题并非与任务相关的信息如何进入网络,而是策略的内部状态是否保留了该信息。基于此视角,我们提出用于机器人运动的SleepWalking(SWAQ),这是一种单阶段端到端框架,利用下一步特权物理重构,在策略学习过程中塑造循环历史表征的保留内容,而部署的执行器仅使用直接的历史到动作路径。在对齐的训练设置下,SWAQ达到比最强的非外感受基线DWAQ高15.0%的峰值平均地形水平,同时每个控制步骤的推理MAC减少44.4%。分层探针进一步显示,与重构物理变量相关的信息可通过策略头线性解码,直至动作输出前的最后一层。补充理论分析将特权变量可恢复性与基于历史的策略类和特权信息策略类之间的可实现回报差距相关联。这些结果表明,语义目标可在无需对部署控制器进行相应架构分解的情况下构建学习过程。

英文摘要

Partially observable locomotion requires a policy to act when task-relevant properties of the robot--environment state are not fully specified by instantaneous observations. Existing approaches often address this challenge by explicitly estimating missing physical variables or processing extended observation histories through structured architectures. We take a different view: partial observability is fundamentally an information-retention problem. The decisive question is not how task-relevant information enters the network, but whether the policy's internal state retains it. Guided by this perspective, we propose SleepWalking for Robot Locomotion (SWAQ), a one-stage end-to-end framework that uses next-step privileged physical reconstruction to shape what a recurrent history representation retains during policy learning, while the deployed actor uses only a direct history-to-action pathway. Under aligned training settings, SWAQ achieves a 15.0\% higher peak mean terrain level than DWAQ, the strongest non-exteroceptive baseline, while using 44.4\% fewer inference MACs per control step. Layerwise probes further show that information associated with the reconstructed physical variables remains linearly decodable through the policy head up to the layer preceding the action output. Complementary theoretical analysis relates privileged-variable recoverability to the achievable-return gap between history-based and privileged-information policy classes. These results suggest that semantic objectives can structure learning without requiring a corresponding architectural decomposition of the deployed controller.

Comments18 pages.13 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑