发表机构
Northwestern Polytechnical University(西北工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出LeCo策略-导纳学习框架,通过多速率反馈和冲突代价引导视觉策略利用固定柔顺性,在四类真实连接器装配任务中实现94%成功率并显著降低接触载荷。
AI 中文摘要
策略学习与柔顺控制为在位姿误差和接触不确定性下实现可靠的自主装配提供了一条有前景的途径。然而,将两者结合并不能确保协调:策略可能持续推压接触面,而控制器却在退让,从而产生持续的载荷而进展有限。为解决这一问题,我们提出LeCo(利用柔顺性),一种策略-导纳学习框架,在固定导纳下通过执行时交互来引导视觉策略。一种多速率反馈机制将高频率的接触交互记录聚合为策略转移奖励。一个集成的冲突代价随后刻画了策略持续加载与控制器卸载之间的对立,而一个方向性的大力尾部代价则捕捉转移内的持续加载事件。与任务完成一起,这些代价鼓励策略利用柔顺性,减少无效加载。我们在四个真实连接器装配任务上评估了LeCo,获得了94%的总体成功率。在各任务中,相对于对比基线,平均成功试验的合力峰值和力矩峰值分别降低了约30%和64%。奖励消融进一步表明,加入冲突塑形可将成功试验中接触条件冲突密度的中位数降低约53%。这些结果支持通过将多速率策略-导纳交互转化为互补的奖励信号,来学习利用固定柔顺性,从而实现有效且低载荷的插装。
英文摘要
Policy learning and compliant control offer a promising route to reliable autonomous assembly under pose errors and contact uncertainty. However, combining them does not ensure coordination: the policy may continue pushing against contact while the controller yields, producing sustained loading with limited progress. To address this problem, we propose LeCo (Leverage Compliance), a policy-admittance learning framework that guides a visual policy through execution-time interaction under fixed admittance. A multirate feedback mechanism aggregates high-rate contact-interaction records into policy-transition rewards. An integrated conflict cost then characterizes sustained policy-loading/controller-unloading opposition, while a directional high-force tail cost captures continued-loading events within a transition. Together with task completion, these costs encourage the policy to leverage compliance with less unproductive loading. We evaluate LeCo on four real connector-assembly tasks, obtaining an aggregate success rate of 94%. Across tasks, mean successful-trial resultant-force and torque peaks decrease by approximately 30% and 64% relative to the comparison baseline. Reward ablation further shows that adding conflict shaping reduces median successful-trial contact-conditioned conflict density by approximately 53%. These results support learning to leverage fixed compliance by turning multirate policy-admittance interaction into complementary reward signals for effective, lower-load insertion.