arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GraspTwin:基于数字孪生的零样本任务导向抓取优化

GraspTwin: Zero-Shot Task-Oriented Grasp Optimization via a Digital Twin

Daniel J. Evans, Yinlong Dai, Simon Stepputtis, Dylan P. Losey

arXiv 2609.30543首次发表:更新:

发表机构

Virginia Tech(弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器人任务导向抓取问题,提出GraspTwin框架,利用数字孪生和基础模型生成语义先验,经贝叶斯优化与物理仿真验证,实现零样本真实世界抓取,成功率提升达33%。

AI 中文摘要

随着机器人从结构化工厂环境过渡到家庭环境,它们需要与日益多样化的物体进行交互。许多任务需要抓取,而仅仅拾起目标物体往往是不够的。考虑一个像“倒咖啡”这样的任务——为了便于后续倒咖啡,机器人应该抓住杯子的把手。现有的基于学习的抓取方法要么寻找与任务基本无关的稳健且无碰撞的抓取(例如,抓住杯子的边缘),要么利用基础模型提出缺乏细粒度物理基础的任务适当抓取位置(例如,伸手去抓把手却错过)。在这项工作中,我们通过一个真实到仿真再到真实的框架来弥合这些方法。基于单个RGB-D观测,我们构建环境的数字孪生,查询大型基础模型以提出与物体功能属性和任务描述一致的抓取,然后优化这些提议以确保鲁棒性和合理性,最后在真实机器人上执行结果。我们的关键见解是,基础模型的抓取提议应被视为语义先验,作为局部无梯度优化的种子。我们利用带有汤普森采样的贝叶斯优化来抽取邻近姿态的批次,随后在域随机化物理滚动下并行评估。最终得到的抓取既具有任务导向性,又物理上可行,可由机器人手臂执行。我们完整的零样本真实世界迁移仅需几分钟,与其他最先进的流程相比,任务导向抓取成功率提高了高达33%。我们的代码可在此处获取:此https URL

英文摘要

As robots transition from structured factory settings into homes, they are required to interact with an ever-increasing variety of objects. Many tasks require grasping, and often it is not sufficient to just pick up the target object. Consider a task like "pouring coffee" --- to facilitate the subsequent pouring, the robot should grasp the mug by its handle. Existing learning-based approaches for grasping either find robust and collision-free grasps that are largely agnostic to the task (e.g., picking up the mug by its rim), or leverage foundation models to propose task-appropriate grasp locations that lack fine-grained physical grounding (e.g., reaching for and missing the handle). In this work, we bridge these approaches with a real-to-sim-to-real framework. Based on a single RGB-D observation, we construct a digital twin of the environment, query a large foundation model to propose grasps that align with the object's affordances and task description, and then optimize the proposals to ensure robustness and plausibility before executing the result on the real robot. Our key insight is that the grasp proposals of the foundation model should be regarded as semantic priors that serve as seeds for local, gradient-free optimization. We leverage Bayesian optimization with Thompson sampling to draw batches of nearby poses, which are subsequently evaluated in parallel under domain-randomized physics rollouts. The resulting grasp is both task-oriented and physically feasible for execution by the robot arm. Our full zero-shot real-world transfer only takes a few minutes and improves task-oriented grasping success by up to 33% as compared to other state-of-the-art pipelines. Our code is available here: https://github.com/VT-Collab/GraspTwin/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑