PartDexTOG: 通过语言驱动的部件分析生成灵活的任务导向抓取
PartDexTOG: Generating Dexterous Task-Oriented Grasping via Language-driven Part Analysis
- College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学)
- College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
PartDexTOG通过语言驱动的部件分析生成灵活的任务导向抓取,提升抓取性能和泛化能力。
AI中文摘要:
任务导向的抓取是机器人操作中的关键但具有挑战性的任务。尽管近年来取得了进展,但现有的大多数方法并未针对灵活的手来解决任务导向的抓取问题。灵活的手提供了更好的精度和通用性,使机器人能够更有效地执行任务导向的抓取。在本文中,我们主张通过提供关于物体功能的详细信息来增强灵活的抓取,通过部件分析。我们提出了PartDexTOG,一种通过语言驱动的部件分析生成灵活的任务导向抓取的方法。输入一个3D物体和一个由语言表示的操作任务,该方法首先通过大语言模型生成与操作任务相关的类别级和部件级抓取描述。然后,开发了一个类别-部件条件扩散模型,根据生成的描述分别生成每个部件的灵活抓取。为了从生成的抓取中选择最合理的组合,我们提出了抓取与部件之间的几何一致性度量。我们证明了我们的方法从大语言模型在物体部件上的开放世界知识推理中受益,这自然促进了对不同几何形状和不同操作任务的抓取生成的学习。我们的方法在OakInk-shape数据集上优于所有先前方法,提高了穿透体积、抓取位移和P-FID,分别提高了3.58%、2.87%和41.43%。值得注意的是,它在处理新类别和任务时表现出良好的泛化能力。
英文摘要:
Task-oriented grasping is a crucial yet challenging task in robotic manipulation. Despite the recent progress, few existing methods address task-oriented grasping with dexterous hands. Dexterous hands provide better precision and versatility, enabling robots to perform task-oriented grasping more effectively. In this paper, we argue that part analysis can enhance dexterous grasping by providing detailed information about the object's functionality. We propose PartDexTOG, a method that generates dexterous task-oriented grasps via language-driven part analysis. Taking a 3D object and a manipulation task represented by language as input, the method first generates the category-level and part-level grasp descriptions w.r.t the manipulation task by LLMs. Then, a category-part conditional diffusion model is developed to generate a dexterous grasp for each part, respectively, based on the generated descriptions. To select the most plausible combination of grasp and corresponding part from the generated ones, we propose a measure of geometric consistency between grasp and part. We show that our method greatly benefits from the open-world knowledge reasoning on object parts by LLMs, which naturally facilitates the learning of grasp generation on objects with different geometry and for different manipulation tasks. Our method ranks top on the OakInk-shape dataset over all previous methods, improving the Penetration Volume, the Grasp Displace, and the P-FID over the state-of-the-art by $3.58\%$, $2.87\%$, and $41.43\%$, respectively. Notably, it demonstrates good generality in handling novel categories and tasks.