arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

条件扩散策略中用于塑造可变形线性物体的在线材料估计

Online Material Estimation for Conditioned Diffusion Policy in Shaping Deformable Linear Objects

Ryunosuke Yamada, Tomohiro Motoda, Yukiyasu Domae, Tokuo Tsuji

arXiv 2609.12634首次发表:更新:

发表机构

Kanazawa University; National Institute of Advanced Industrial Science and Technology (AIST)(金泽大学; 产业技术综合研究所(AIST))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对可变形线性物体形状控制中材料属性导致策略泛化难的问题,提出在线估计材料标签并条件化扩散策略的方法,在真实机器人任务中达到与已知材料标签相当的60.8%成功率。

AI 中文摘要

可变形线性物体(DLOs)的形状控制对于模仿学习而言具有挑战性,因为变形行为会随刚度、弹性等材料属性而变化,因此即使目标形状相同,单一策略也必须针对不同物体生成不同的动作序列。我们提出了一种以材料标签为条件的扩散策略,该材料标签在操作过程中在线估计。一个循环估计网络从多视角图像和机器人关节状态的时间序列中预测被抓取物体的材料标签,并在每次推理步骤中将预测的标签作为扩散策略的条件。我们收集了覆盖四种DLO材料和三种凹槽放置任务的480次真实机器人演示,并比较了每种材料的专家策略、无材料标签的任务条件策略、以真实材料标签为条件的策略以及所提出的策略。以真实材料标签为条件相比仅任务条件策略,将平均成功率从45.8%提升至60.0%,而所提出的策略在没有任何先验材料信息的情况下达到了60.8%,与给定真实标签的策略相当。事后分析表明,估计器从操作观测中提取了与材料相关的信息,扩散策略对由此产生的条件信号作出响应,而一个显著的失败案例与两种相似材料之间的持续混淆有关。

英文摘要

Shape control of deformable linear objects (DLOs) is challenging for imitation learning because deformation behavior varies with material properties such as stiffness and elasticity, so a single policy must generate different action sequences for different objects even when the goal shape is identical. We propose a diffusion policy conditioned on material labels that are estimated online during manipulation. A recurrent estimation network predicts the material label of the grasped object from the time series of multi-view images and robot joint states, and the predicted label conditions the diffusion policy at every inference step. We collected 480 real-robot demonstrations covering four DLO materials and three groove-placement tasks, and compared per-material specialist policies, a task-conditioned policy without material labels, a policy conditioned on ground-truth material labels, and the proposed policy. Conditioning on ground-truth material labels improved the average success rate from 45.8% to 60.0% over the task-only policy, and the proposed policy reached 60.8% without any prior material information, matching the policy given ground-truth labels. A post-hoc analysis shows that the estimator extracts material-related information from the manipulation observations and that the diffusion policy responds to the resulting conditioning signal, while the one pronounced failure case is associated with persistent confusion between two similar materials.

Comments10 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑