HiWE:具有视觉关键点增强的层次化世界知识模型用于零样本3D路径规划
HiWE: Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning
浏览论文内容
中文总结 AI 辅助
HiWE提出一种结合视觉关键点增强的层次化世界知识模型,通过点基接口连接视觉接地与语言规划,实现零样本3D路径规划,并在模拟和物理任务上验证有效性。
中文摘要 AI 辅助
机器人演示生成需要一个系统来识别交互应发生的位置、规划可行的运动并执行所需的接触。HiWE通过视觉接地与基于语言的规划之间的点基接口连接这些决策。PointVLM经过指令微调,利用点标注、分割衍生样本、机器人观测和视觉问答数据的混合,将任务相关对象与图像坐标关联起来。深度测量将这些预测提升为语义3D表示。语言规划器3DLLM利用该表示来指定末端执行器路点和夹爪命令,而混合抓取模块则解决局部抓取姿态。评估涵盖14个模拟操作任务和4个物理机器人任务,并包括对视觉训练数据、空间输入和抓取选择的消融实验。这里,零样本执行指的是在没有任务特定演示训练的情况下进行部署;视觉模型在微调期间使用现有的机器人数据。本文描述了该框架的原始点基公式;其与后续GeneralVLA扩展的关系在引言中详述。
英文摘要
Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-relevant objects with image coordinates using a mixture of point annotations, segmentation-derived samples, robot observations, and visual question answering data. Depth measurements lift these predictions into a semantic 3D representation. A language planner, 3DLLM, uses this representation to specify end-effector waypoints and gripper commands, while a hybrid grasping module resolves local grasp poses. The evaluation covers 14 simulated manipulation tasks and four physical-robot tasks, together with ablations of the visual training data, spatial inputs, and grasp selection. Here, zero-shot execution refers to deployment without task-specific demonstration training; the visual model uses existing robot data during fine-tuning. This paper describes the original point-based formulation of the framework; its relationship to the subsequent GeneralVLA extension is detailed in the introduction.
发表机构
- University of Chinese Academy of Sciences(中国科学院大学)
- State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, CAS(中国科学院脑认知与脑启发智能技术国家重点实验室)
- The University of Hong Kong(香港大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。