发表机构
Anker Spatial Perception Lab(安克空间感知实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究有腿机器人在动态环境中导航问题,核心方法是用辅助预测器和正则化增强反应式导航主干进行预测训练,主要贡献是提高导航成功率、降低碰撞率并实现零样本模拟到现实的迁移。
AI 中文摘要
用于有腿机器人的强化学习导航策略从当前观察和短期记忆中被动选择动作,预测移动障碍物未来演变的能力有限。在动态环境中,这种被动性导致机器人反应过晚。训练期间对策略循环状态应用轻量级预测监督可编码预期障碍物动态。我们用辅助JEPA风格预测器和SIGReg正则化增强反应式LSTM - SRU导航主干。在模拟和现实世界动态障碍物导航基准测试中,该方法大幅提高导航成功率并降低碰撞率,在Unitree Go2上实现零样本模拟到现实的迁移。
英文摘要
Reinforcement-learning navigation policies for legged robots select actions reactively from current observations and short-term memory, with limited capacity to anticipate how moving obstacles will evolve in the near future. In dynamic environments, this reactivity causes the robot to respond too late because collision risk depends on short-horizon scene structure rather than on current obstacle positions alone. Lightweight predictive supervision applied to the policy's recurrent state during training can encode anticipatory obstacle dynamics without modifying the inference-time controller. We augment a reactive LSTM-SRU navigation backbone with an auxiliary JEPA-style predictor and SIGReg regularization: during training, the predictor supervises the deterministic hidden state to anticipate its own next state; at inference, it is fully discarded, incurring zero additional computational cost. On simulated and real-world navigation benchmarks with dynamic obstacles, our method substantially improves navigation success while reducing collision rates through the predictive training signal alone, without additional inference-time parameters. Real-robot deployment on a Unitree Go2 demonstrates zero-shot sim-to-real transfer: the controller navigates cluttered indoor and dynamic outdoor environments without fine-tuning, with evasive behavior consistent with the collision reduction observed in simulation.
Comments10 pages, 9 figures