GraphPoint:面向组合式机器人操作任务的语义实体图与点轨迹方法
GraphPoint: Semantic Entity Graphs and Point Trajectories for Compositional Robot Manipulation
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出GraphPoint方法,通过语义实体图与点轨迹预测实现组合式机器人操作,并在CoMani基准上验证了子任务内与跨子任务的语言指令泛化能力。
AI中文摘要:
机器人操作策略往往难以泛化到其演示之外的情境,即使新指令涉及熟悉的物体和行为也是如此。当语言和场景在训练过程中高度相关时,策略可能学习到固定的视觉-动作映射,而非响应所请求的行为。我们在两个层面研究组合式复用:在子任务内部,组合熟悉的实体、动作类型和动作修饰符;在子任务之间,在未见过的长时程任务中复用已学习的子任务。我们引入了CoMani基准,该基准通过受控的数据划分来评估这两种能力。匹配的初始场景和对单一语义因子的受控修改,鼓励策略依赖语言而非视觉捷径。我们进一步提出了GraphPoint方法,该方法通过预测未来的夹爪点轨迹,并利用机器人几何模型将其转换为动作,从而将语义实体图与几何控制相连接。该框架根据语义角色组织夹爪和物体,并根据动作类型和修饰符调节其交互,同时预测的进度在执行过程中指导状态转换。在CoMani上的实验和消融研究验证了我们的方法在指令相关泛化方面两个层面的有效性。代码将在GraphPoint发布。
英文摘要:
Robot manipulation policies often struggle to generalize beyond their demonstrations, even when new instructions involve familiar objects and behaviors. When language and scenes are strongly correlated during training, a policy can learn a fixed visual-action mapping rather than respond to the requested behavior. We investigate compositional reuse at two levels: within a subtask, combining familiar entities, action types, and action modifiers; and across subtasks, reusing learned subtasks in unseen long-horizon tasks. We introduce CoMani, a benchmark with controlled splits for evaluating both capabilities. Matched initial scenes and controlled changes to a single semantic factor encourage reliance on language rather than visual shortcuts. We further propose GraphPoint, which connects semantic entity graphs to geometric control by predicting future gripper point trajectories and converting them into actions using robot geometry. The framework organizes the gripper and objects by semantic roles and conditions their interactions on action types and modifiers, while predicted progress guides transitions during execution. Experiments and ablations on CoMani validate the effectiveness of our method for instruction-dependent generalization at both levels. Code will be released at GraphPoint.