约束锚定强化学习用于接触丰富机器人插入中的可变阻抗控制
Constraint-Grounded Reinforcement Learning for Variable Impedance Control in Contact-Rich Robotic Insertion
浏览论文内容
中文总结 AI 辅助
针对机器人插入中固定增益需手动调整的问题,提出约束锚定强化学习框架,在力限制下在线调整阻抗增益,实验显示成功率显著优于固定增益基线。
中文摘要 AI 辅助
在不确定接触下的机器人插入任务中,轴向力限制和合适的控制器增益因任务而异。因此,单一的固定增益在不同任务条件下难以保持适用,使得传统阻抗控制器依赖手动重新调整。为消除手动重新调整,我们提出了约束锚定强化学习(CG-RL),一种用于在线增益自适应的可变阻抗框架。以力限制和接触反馈为条件,策略输出残余运动、插入速率和请求增益。控制器将该增益投影到允许范围内,而不将该范围暴露给策略。这种分离使得单一策略能够在不同力限制下运行,无需重新训练或手动调整。我们在模拟的斜向插入任务上,使用五个训练种子评估CG-RL。CG-RL实现了$85.8\pm7.7\\%$(均值$\pm$标准差)的插入成功率,且未违反力限制,同时保持施加增益在允许范围内。作为对比,使用中点增益的固定增益基线实现了$50.1\\%$的成功率。该策略根据指定的力限制连续调整其插入速率,并进一步泛化到训练范围以上的更宽松力限制。相比之下,没有力限制输入的相同actor不表现出这种自适应。施加增益被保证保持在允许范围内,而力限制的满足是经验验证而非正式保证。
英文摘要
In robotic insertion under uncertain contact, the axial force limit and the appropriate controller gain vary across tasks. As a result, a single fixed gain is unlikely to remain suitable across different task conditions, making conventional impedance controllers reliant on manual retuning. To eliminate manual retuning, we propose Constraint-Grounded Reinforcement Learning (CG-RL), a variable impedance framework for online gain adaptation. Conditioned on the force limit and contact feedback, the policy outputs a residual motion, an insertion rate, and a requested gain. The controller projects this gain into the admissible range without exposing the range itself to the policy. This separation allows a single policy to operate under different force limits without retraining or manual retuning. We evaluate CG-RL on simulated oblique insertion across five training seeds. CG-RL achieves an $85.8\pm7.7\%$ (mean $\pm$ SD) success rate of insertions without violating the force limit, while keeping the applied gain within the admissible range. As a comparison, a fixed-gain baseline using the midpoint gain achieves a success rate of $50.1\%$. The policy adapts its insertion rate continuously to the specified force limit and further generalizes to more permissive force limits above the training range. In contrast, the same actor without force-limit input does not exhibit this adaptation. The applied gain is guaranteed to remain within the admissible range, while force-limit satisfaction is validated empirically rather than guaranteed formally.
发表机构
- The University of Tennessee(田纳西大学)
机构由 AI 辅助整理,请以论文原文为准。