发表机构
Neuromeka Co., Ltd.(纽若美卡有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对接触丰富操作中传统控制器与局部约束不匹配的问题,提出在仿真中学习本体感觉反射策略,作为高层控制器下的冻结执行层,在箱体提升、销钉插入和表面跟随中提升力控制与成功率。
AI 中文摘要
在接触丰富的操作过程中,机器人与环境之间的相互作用携带了关于局部几何形状的信息:表面阻止穿透,孔引导销钉。利用这些相互作用的控制器可以在保持任务意图的同时顺应环境约束。机器人学习系统通常使用位置控制、混合力-位置控制或笛卡尔阻抗控制。它们规定的跟踪目标、刚度或力控制方向可能与局部约束不匹配,从而降低性能。我们在仿真中基于三种简单的交互原语(弹簧、平面和导轨)学习了一种本体感觉反射策略。该策略利用状态历史将任务空间命令映射到关节位置目标,无需直接的力或几何测量。训练完成后,它作为冻结的执行层位于高层控制器之下。我们在双臂箱体提升、销钉插入和表面跟随任务中进行了评估。在箱体提升中,反射策略将力保持在阈值以下,而基线方法失败。在粗糙表面跟随中,它保持在10 N参考值以下,与调优的混合力-位置控制相当。在0.02毫米销钉插入中,它将平均估计接触力减少到基线的一半以下,同时将硬件成功率从至多22%提高到36-58%。通过将接触响应与命令生成分离,反射策略为运动规划器、学习策略和远程操作员提供了稳健的接触丰富执行能力。
英文摘要
During contact-rich manipulation, interactions between a robot and its environment carry information about local geometry: a surface prevents penetration, a bore guides a peg. A controller that exploits these interactions can comply with environmental constraints while preserving task intent. Robot-earning systems commonly use position, hybrid force-position, or Cartesian impedance control. Their prescribed tracking objectives, stiffness, or force-control directions may not match local constraints and may degrade performance. We learn a proprioceptive reflex policy in simulation on three simple interaction primitives: a spring, a plane, and a rail. The policy maps task-space commands to joint-position targets using state history, without direct force or geometrical measurements. Once trained, it serves as a frozen execution layer beneath higher-level controllers. We evaluate it in dual-arm box lifting, peg insertion, and surface following. In box lifting, the reflex kept the force below the threshold while the baseline failed. In rough-surface following it stayed below the 10 N reference, on par with tuned hybrid force-position control. In 0.02 mm peg insertion It reduced mean estimated contact force to less than half that of the baseline while increasing hardware success rates from at most 22 % to 36-58 %. By separating contact response from command generation, the reflex policy provides motion planners, learned policies, and teleoperators with robust contact-rich execution.