arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

势场动作表示用于接触丰富操作中的强化学习

Potential-Field Action Representation for Reinforcement Learning in Contact-Rich Manipulation

Xinyu Liu, Gökhan Solak, Arash Ajoudani

arXiv 2609.21609首次发表:更新:

发表机构

Istituto Italiano di Tecnologia; Università di Genova(意大利理工学院; 热那亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出PA-RL框架,以人工势场作为动作表示,通过调整势场参数生成引导方向,在插轴入孔任务中实现100%仿真成功率,并减少运动变化,且可迁移至真实机器人。

AI 中文摘要

无模型强化学习可以通过试错交互获取接触丰富的机器人操作技能,但通常要求策略同时学习任务策略和低级运动生成。在此设定下,动作表示至关重要,因为它决定了策略输出如何转换为机器人运动,从而影响探索和物理执行。直接笛卡尔命令接口要求策略在每个决策步骤生成运动,将任务级适应与连续低级控制耦合,增加了学习负担。我们提出PA-RL,一种使用人工势场作为动作表示的强化学习框架。策略不直接命令运动,而是调整类能量势场的参数,该势场生成状态相关的引导方向,并通过笛卡尔阻抗控制器执行。我们在插轴入孔任务上评估PA-RL,该任务具有非线性动力学和不连续接触转换,是代表性的接触丰富任务。在仿真中,PA-RL与使用相同RL算法的笛卡尔速度、笛卡尔位姿和可变阻抗动作空间进行比较。PA-RL是唯一在分配的训练时间内达到100%评估成功率的方法,而最佳基线达到92.6%。相对于最佳基线,PA-RL还将关节扭矩变化减少55.4%,笛卡尔加速度变化减少70.8%,且无需在奖励中显式添加运动质量惩罚。仿真训练的策略进一步在无需微调的情况下完成9/9次真实机器人插入,展示了所学势场接口的部署可行性。

英文摘要

Model-free reinforcement learning can acquire contact-rich robotic manipulation skills through trial-and-error interaction, but it often requires the policy to learn both task strategy and low-level motion generation. In this setting, the action representation is critical because it determines how policy outputs are converted into robot motion, shaping both exploration and physical execution. Direct Cartesian command interfaces require the policy to generate motion at every decision step, coupling task-level adaptation with continuous low-level control and increasing the learning burden. We propose PA-RL, a reinforcement-learning framework that uses artificial potential fields as the action representation. Instead of commanding motion directly, the policy adapts the parameters of an energy-like potential field, which generates a state-dependent guidance direction executed through a Cartesian impedance controller. We evaluate PA-RL on peg-in-hole insertion, a representative contact-rich task with nonlinear dynamics and discontinuous contact transitions. In simulation, PA-RL is compared with Cartesian velocity, Cartesian pose, and variable-impedance action spaces using the same RL algorithm. PA-RL is the only method to reach a 100% evaluation success rate within the allotted training time, while the best baseline reaches 92.6%. It also reduces joint-torque variation by 55.4% and Cartesian acceleration variation by 70.8% relative to the best baseline, without explicit motion-quality penalties in the reward. The simulation-trained policy further completes 9/9 real-robot insertions without fine-tuning, demonstrating the deployment feasibility of the learned potential-field interface.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑