arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PAKT:用于强化学习的物理对齐动觉示教

PAKT: Physically-Aligned Kinesthetic Teaching for Reinforcement Learning

Lars Johannsmeier, Yashraj Narang

arXiv 2609.25630首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PAKT提出一种物理对齐的动觉示教框架,通过导纳控制与高性能控制栈,在接触丰富的工业操作中显著降低循环时间和干预次数。

AI 中文摘要

现实世界的强化学习(RL)系统仍然难以应对接触丰富的工业操作需求,包括微米级精度、超过99%的成功率以及人类水平的循环时间。尽管离策略算法可以通过利用演示和干预来提高性能,但一个关键瓶颈是缺乏一种直观的界面来收集此类指导,同时满足物理系统和策略的约束。我们提出了PAKT,一个用于RL中动觉示教的框架。与遥操作方法不同,PAKT依赖于工业中广泛使用的动觉引导。然而,动觉引导的一个关键弱点是操作员可能沿着机器人或策略无法物理复现的轨迹(例如速度、加速度、加加速度)移动机器人。使用PAKT,操作员通过导纳控制引导机器人,该控制将人类施加的力映射为运动。下游参考生成器应用与策略执行期间相同的运动学限制,使收集的轨迹保持在这些限制内。为了支持这一示教界面并配备适当的执行层,PAKT增加了一个高性能控制栈,将低频RL动作映射为高频扭矩指令。它由一个参考生成器和随后的阻抗控制器组成,其中参考生成器在改善接触处理和产生更平滑的策略动作的同时,保持了阻抗控制器的跟踪性能。在四个插入和工业装配基准(包括数据中心计算托盘)的报告运行中,相对于HIL-SERL基线,端到端系统将循环时间减少了23%-48%,累计干预次数减少了62%-86%。项目网站:此https URL }{ 此https URL

英文摘要

Real-world reinforcement learning (RL) systems still struggle with the demands of contact-rich industrial manipulation, including micrometer-level precision, success rates above 99%, and human-level cycle times. Although off-policy algorithms can improve performance by leveraging demonstrations and interventions, a key bottleneck is the lack of an intuitive interface for collecting such guidance while complying with constraints of the physical system and the policy. We propose PAKT, a framework for kinesthetic teaching in RL. As opposed to teleoperation approaches, PAKT relies on kinesthetic guidance, which is widely used in industry. However, a critical weakness of kinesthetic guidance is the possibility for the operator to move the robot along trajectories (e.g., velocities, accelerations, jerk) that the robot and/or policy cannot physically reproduce. Using PAKT, operators guide the robot through admittance control, which maps human-applied forces to motion. The downstream reference generator applies the same kinematic limits used during policy execution, keeping the collected trajectories within these limits. To support this teaching interface with an appropriate execution layer, PAKT adds a high-performance control stack that maps low-frequency RL actions to high-frequency torque commands. It consists of a reference generator and subsequent impedance controller, where the reference generator preserves the tracking performance of the impedance controller while improving contact handling and producing smoother policy actions. Across the reported runs on four insertion and industrial assembly benchmarks, including a data center compute tray, the end-to-end system reduces cycle time by 23%-48% and cumulative intervention count by 62%-86% relative to the HIL-SERL baseline. Project website: https://pakt-website.github.io/pakt-website}{https://pakt-website.github.io/pakt-website

Comments17 pages, 6 figures, 9 tables, Conference on Robot Learning

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑