arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知识引导的分层策略学习用于紧公差下高精度圆柱装配

Knowledge-Guided Hierarchical Policy Learning for High-Precision Cylindrical Assembly under Tight Tolerances

Binbin Lian, Xinyu Liu, Tao Sun

arXiv 2609.06522首次发表:更新:

发表机构

Tianjin University(天津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出知识引导的分层策略学习框架,结合行为克隆与TD3算法,实现高精度圆柱装配,在仿真和真实环境中均表现出高学习效率、强适应性和稳定性。

AI 中文摘要

提出了一种混合分层学习框架,以实现直径为170毫米、公差为0.1毫米的圆柱形部件的高精度装配。下层网络通过行为克隆(BC)整合专家经验,赋予机器人类似人类的直觉,并结合双延迟深度确定性策略梯度(TD3)算法以增强训练稳定性和鲁棒性。上层网络基于启发式规则动态调整下层决策,确保操作的灵活性。构建了一个仿真模型,在迁移到真实世界之前进行学习,从而允许高效且安全的训练。比较结果表明,奖励曲线在500个回合内收敛,表明学习效率高。该方法还展现出对初始条件和位姿误差更好的适应性,即使在极端条件下也能获得满意的成功率。此外,该方法在高斯噪声干扰下表现出良好的稳定性。在真实世界中,圆柱段的装配轨迹显示出更平滑的运动和更小的波动。

英文摘要

A hybrid hierarchical learning framework is proposed to achieve high-precision assembly of 170mm cylindrical components with tolerance of 0.1mm. The lower-level network integrates expert experience through Behavior Cloning (BC), giving the robot human-like intuition, and incorporates the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm to enhance training stability and robustness. The upper-level network dynamically adjusts the lower-level decisions based on heuristic rules, ensuring flexibility in operations. A simulated model is constructed to learn before transferring to real world. An efficient and safe training is allowed. Comparisons show that the reward curve converges within 500 episodes, indicating high learning efficiency. It also demonstrates better adaptability to initial conditions and pose errors, achieving satisfactory success rates even under extreme conditions. Moreover, the method exhibits good stability under Gaussian noise interference. In the real world, the assembly trajectory of the cylindrical segment shows smoother motion and less fluctuation.

Comments25 pages, 14 figures. Submitted to Robotics and Computer-Integrated Manufacturing (Elsevier)

DOI:10.1016/j.rcim.2026.103378

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑