arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31112cs.RO

DualManip:通过双路径语义推理与几何自适应的智能体动态操作

DualManip: Agentic Dynamic Manipulation via Dual-Path Semantic Reasoning and Geometric Adaptation

Chengxi Li, Yan Di, Yingyue Li, Ruida Zhang, Mingyang Li, Xiangyang Ji

首次发表
浏览论文内容

中文总结 AI 辅助

DualManip通过双路径框架解耦语义推理与几何自适应,实现动态场景下鲁棒且高效的机器人操作,几何自适应速度提升约46倍。

中文摘要 AI 辅助

视觉语言模型(VLMs)能够实现机器人操作的开放词汇推理,但其高推理延迟限制了在动态场景中的响应能力。然而,许多场景变化虽然改变了物体几何形状,却并未使任务意图失效。我们提出了DualManip,一种双路径框架,将低频的语义推理与响应式的几何自适应解耦。语义路径对任务进行分解并确定与任务相关的交互,随后通过约束求解模块进行位姿优化。在执行过程中,几何路径通过形状自适应网络,从实时RGB-D观测中持续更新模板与观测之间的对应关系。这些对应关系将任务相关的抓取接触点在不同观测间传递,从而在物体运动和非刚性变形下实现在线抓取重建。信息交互模块通过从语义定位初始化任务相关抓取、验证几何更新,并在更新失败时触发语义重规划,来连接两条路径。真实世界评估涵盖六个操作任务,包括非刚性变形、铰接式重构、刚性运动和精密装配,并在三种设置下进行:静态、单次变化和连续动态。DualManip展现出优越的操作鲁棒性,尤其在连续场景变化下,其几何自适应速度比智能体验证和语义重规划快约46倍。我们的项目页面:此https URL。

英文摘要

Vision-language models (VLMs) enable open-vocabulary reasoning for robot manipulation, but their high inference latency limits responsiveness in dynamic scenes. Many scene changes, however, alter object geometry without invalidating task intent. We present DualManip, a dual-path framework that decouples infrequent semantic reasoning from responsive geometric adaptation. The semantic path decomposes the task and grounds task-relevant interactions, followed by a constraint-solving module for pose optimization. During execution, the geometric path continuously updates template-to-observation correspondences from live RGB-D observations via a shape-adaptive network. These correspondences transfer task-relevant grasp contacts across observations, enabling online grasp reconstruction under object motion and non-rigid deformation. The Information Interaction Module bridges the two paths by initializing task-relevant grasps from semantic grounding, validating geometric updates, and triggering semantic replanning upon update failures. Real-world evaluation spans six manipulation tasks covering non-rigid deformation, articulated reconfiguration, rigid motion, and high-precision assembly across three settings: static, single-change, and continuous dynamic. DualManip demonstrates superior manipulation robustness, particularly under continuous scene changes, while achieving geometric adaptation approximately 46$\times$ faster than agentic verification and semantic replanning. Our project page: https://lichengxi1.github.io/Dualmanip.

发表机构

  • Tsinghua University(清华大学)
  • Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
  • Beijing Institute of Control Engineering(北京控制工程研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑