发表机构
Southwest Jiaotong University; University of Leeds; Southeast University; University of Oxford; University of Surrey(西南交通大学; 利兹大学; 东南大学; 牛津大学; 萨里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RodForesight通过视觉伺服和世界模型增强的扩散策略,分两阶段解决细长杆插入问题,将成功率从88.9%提升至96.7%。
AI 中文摘要
细长杆插入出现在精密制造中,其中毫米级直径和紧密间隙要求精确的感知和控制。传统的销孔装配方法假设物体是刚性的,其尖端位姿相对于夹持器固定。这一假设对于高长径比的杆不再成立,因为杆在操作过程中可能弯曲,使其尖端运动取决于杆的构型、抓取方式、材料特性和接触情况。我们提出了RodForesight,一个将任务分解为两个阶段的学习框架:1)粗略接近,使用视觉伺服将各种初始构型映射到一个紧凑的近孔交接区域;2)预测性插入,执行精细对准并完成插入。值得注意的是,这两个阶段可以封装成一个端到端的设计。在插入过程中,扩散策略生成候选动作块,而一个动作条件世界模型预测它们对杆孔对准的影响。这种执行前评估使RodForesight能够在执行前基于预测的倾斜和径向误差选择最佳动作块。实验研究了不同阶段和端到端设置的性能,其中RodForesight将成功率从88.9%提高到96.7%,相比扩散策略等基线方法。
英文摘要
Slender rod insertion arises in precision manufacturing, where millimetre scale diameter and tight clearances demand accurate perception and control. Conventional peg-in-hole methods assume a rigid object whose tip pose is fixed relative to the gripper. This assumption breaks down for a high aspect ratio rod, which can bend during manipulation, making its tip motion dependent on the rod configuration, grasp, material properties, and contact. We present RodForesight, a learning framework that factorises the task into two stages: 1) coarse approaching, which uses visual servoing to map diverse initial configurations into a compact near hole hand-off region; and 2) predictive insertion, which performs fine alignment and completes the insertion. It is worth noting that the two stages can be wrapped into an end-to-end design. During insertion, a diffusion policy generates candidate action chunks, while an action conditioned world model predicts their effects on rod-hole alignment. This pre-execution evaluation enables RodForesight to select the best action chunk based on predicted tilt and radial errors before execution. Experiments investigate the performance of different stages and the end-to-end setting, where RodForesight improves the success rate from 88.9% to 96.7%, compared to baseline methods such as diffusion policy.