ZeroDexGrasp:基于提示多阶段语义推理的零样本任务导向灵巧抓取合成
ZeroDexGrasp: Zero-Shot Task-Oriented Dexterous Grasp Synthesis with Prompt-Based Multi-Stage Semantic Reasoning
- Shenzhen University(深圳大学)
- Sun Yat-sen University(中山大学)
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有任务导向灵巧抓取方法依赖标注数据、泛化性差的问题,提出ZeroDexGrasp框架,通过提示式多阶段语义推理与接触引导优化实现零样本高质量灵巧抓取。
AI中文摘要:
任务导向灵巧抓取在机器人操作与人机交互领域具备广阔应用前景,但现有多数方法严重依赖高成本标注数据实现特定任务的语义对齐,难以泛化至多样化物体与任务指令场景。本研究提出ZeroDexGrasp,这是一种融合多模态大语言模型与抓取优化的零样本任务导向灵巧抓取合成框架,可生成与特定任务目标、物体可供性精准对齐的类人抓取位姿。具体而言,ZeroDexGrasp采用基于提示的多阶段语义推理,从任务与物体语义中推断初始抓取配置与物体接触信息,再通过接触引导的抓取优化对位姿进行修正,保障物理可行性与任务对齐度。实验结果表明,ZeroDexGrasp可在多种未见过的物体类别与复杂任务需求下实现高质量零样本灵巧抓取,推动机器人抓取向更强泛化性、更高智能化方向发展。
英文摘要:
Task-oriented dexterous grasping holds broad application prospects in robotic manipulation and human-object interaction. However, most existing methods still struggle to generalize across diverse objects and task instructions, as they heavily rely on costly labeled data to ensure task-specific semantic alignment. In this study, we propose \textbf{ZeroDexGrasp}, a zero-shot task-oriented dexterous grasp synthesis framework integrating Multimodal Large Language Models with grasp refinement to generate human-like grasp poses that are well aligned with specific task objectives and object affordances. Specifically, ZeroDexGrasp employs prompt-based multi-stage semantic reasoning to infer initial grasp configurations and object contact information from task and object semantics, then exploits contact-guided grasp optimization to refine these poses for physical feasibility and task alignment. Experimental results demonstrate that ZeroDexGrasp enables high-quality zero-shot dexterous grasping on diverse unseen object categories and complex task requirements, advancing toward more generalizable and intelligent robotic grasping.