arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

非夹持投掷:强化学习视角

Non-Prehensile Throwing: A Reinforcement Learning Perspective

Abdullah Mustafa, Ryo Hanai, Ixchel G. Ramirez-Alpizar, Floris Erich, Ryoichi Nakajo, Yukiyasu Domae, Tetsuya Ogata

arXiv 2609.00771首次发表:更新:

发表机构

National Institute of Advanced Industrial Science and Technology (AIST)(日本国立先进工业科学技术研究院(AIST))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对非夹持式机器人投掷,提出基于强化学习的方法,结合滑动与滚动接触模式,在仿真及真实UR5e机器人上实现高投掷成功率,可处理大型、沉重物体。

AI 中文摘要

机器人投掷可实现快速物体搬运,拓展机器人可达工作空间,超越传统的抓取-放置操作。夹持式(基于抓取)投掷适用于可抓取物品,而非夹持式(无抓取)投掷更适合大型、沉重和/或可变形物体。现有方法依赖基于模型的优化,采用简化的接触模型(如动态抓取)和低维轨迹参数化,这限制了解决方案质量和可达工作空间。我们提出一种强化学习方法,该方法额外利用滑动和滚动接触模式,无需解析接触模型或自定义参数化,直接优化关节空间轨迹。马尔可夫决策过程(MDP)被构建为一个动力学系统,其根据投掷目标、物体模型和初始配置演化机器人的关节状态。关节冲击轨迹以低控制率离线规划,并上采样为平滑的高率速度命令用于部署。为实现仿真到真实的迁移,我们通过最小冲击系统辨识最小化机器人动力学差距,并训练感知不确定性的策略以减轻物体建模误差,尤其是对动态摩擦的敏感性。在仿真中,该策略在数千种配置下达到99%的成功率,并能泛化到未见过的物体。敏感性分析显示其对质量不确定性具有鲁棒性,但对动态摩擦高度敏感,这与基于滑动的释放机制一致。在UR5e机器人(以接近其物理极限的5 m/s末端执行器速度运行)上零样本部署,我们的方法可将包括790克重物和20×20×28厘米大尺寸物体在内的各类物体投掷至最远350厘米距离或最高180厘米高度,在真实世界中达到97%的成功率。

英文摘要

Robotic throwing enables fast object transport and extends a robot's reachable workspace beyond traditional pick-and-place. While prehensile (grasp-based) throwing works well for graspable items, non-prehensile (grasp-free) throwing is better suited for large, heavy, and/or deformable objects. Existing approaches rely on model-based optimization with simplified contact models (e.g., dynamic grasping) and low-dimensional trajectory parameterizations, which limit solution quality and reachable workspace. We propose a reinforcement learning approach that additionally leverages sliding and rolling contact modes and directly optimizes joint-space trajectories without analytical contact models or custom parameterizations. The Markov Decision Process (MDP) is formulated as a dynamical system that evolves the robot's joint state conditioned on the throwing target, object model, and initial configuration. Joint-jerk trajectories are planned offline at a low control rate and upsampled into smooth, high-rate velocity commands for deployment. For sim-to-real transfer, we minimize the robot-dynamics gap through minimum-jerk system identification and train uncertainty-aware policies to mitigate object-modeling errors, particularly sensitivity to dynamic friction. In simulation, the policy achieves 99% success across thousands of configurations and generalizes to unseen objects. Sensitivity analysis shows robustness to mass uncertainty but high sensitivity to dynamic friction, consistent with the sliding-based release mechanism. Deployed zero-shot on a UR5e operating near its physical limits (5 m/s end-effector velocity), our method throws diverse objects including heavy (790 g) and large (20x20x28 cm) items to targets up to 350 cm distance or 180 cm elevation, achieving a 97% real-world success rate.

Comments8 pages, 9 figures, Accepted to IEEE IROS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑