arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

时空敏捷性:面向视觉引导的动态四足机器人拦截的时间约束强化学习

Spatiotemporal Agility: Time-Constrained Reinforcement Learning for Vision-Guided Dynamic Quadrupedal Interception

Yidong Zhu, Zibo Dai, Tongning Zhang, Leixin Chang, Hua Chen

arXiv 2608.06907首次发表:更新:

发表机构

Zhejiang University; LimX Dynamics Technology Co., Ltd.(浙江大学; 灵蜥动力科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对四足机器人动态接球任务,提出结合视觉模块与直接目标条件RL策略的集成框架,构建实时闭环拦截系统,提升了接球成功率并减小了sim-to-real性能差距。

AI 中文摘要

腿式机器人需要具备强大的敏捷性,以在受限时间内感知复杂动态环境并与之交互。然而,现有大多数四足机器人运动研究依赖于速度跟踪策略,难以在严格时间约束下到达精确目标;此外,由于传感器延迟和处理延迟,将实时感知与敏捷运动结合以应对高度动态目标仍具挑战性。为具体研究和基准测试动态场景下的此类敏捷性,本文提出一项针对腿式机器人的极具挑战性的接球任务。本文提出一种集成框架,结合用于着陆点与时间预测的视觉模块,以及直接以位置和时间为条件的强化学习(RL)运动策略,而非中间速度指令。除方法设计外,本研究还提供一项系统层面的贡献,即构建完整的实时机器人拦截系统,将多摄像头感知、在线轨迹预测、低延迟目标通信以及仿真到真实(sim-to-real)运动控制整合为闭环部署流水线。通过显式预测未来时空目标,本方法可缓解动态拦截过程中的感知延迟。我们针对腿式机器人开展了大量接球实验,与速度跟踪基线的对比实验表明,本研究提出的直接目标条件方法,在捕获预测着陆点在2米范围内、飞行时间在0.8至1.2秒之间的球时,成功率更高,证明机器人在测试设置下成功完成了动态接球任务;此外,本策略在部署后表现出更小的性能差距,表明这些实验中仿真到真实的行为得到了改善。

英文摘要

Legged robots require robust agility to perceive and interact with complex and dynamic environments within a constrained time. However, most existing quadruped locomotion works rely on velocity-tracking policy, which struggle to reach precise targets within strict temporal constraints. Moreover, integrating real-time perception with agile locomotion for highly dynamic targets remains challenging due to sensor latency and processing delays. To concretely study and benchmark such agility in dynamic settings, we introduce a challenging ball-catching task for legged robots. This paper proposes an integrated framework that combines a vision module for landing point and time prediction with a direct position and time conditioned RL locomotion policy, instead of intermediate velocity commands. Beyond the method design, this work presents a system-level contribution that completes real-time robotic interception system that integrates multi-camera perception, online trajectory prediction, low-latency target communication, and sim-to-real locomotion control into a closed-loop deployment pipeline. By explicitly predicting the future spatial-temporal target, our approach mitigates perception latency during dynamic interception. We conducted extensive ball-catching experiments for the legged robot. Through comparative experiments against a velocity-tracking baseline, our direct target-conditioned approach achieves a higher success rate in catching balls with predicted landing spots within 2 meters and flight times between 0.8 and 1.2 seconds. This shows that the robot has successfully completed the dynamic ball-catching task under our tested setup. Furthermore, our policy exhibits a smaller performance gap after deployment, suggesting improved sim-to-real behavior in these trials.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑