发表机构
Nanjing University; Ant Group; Zhejiang University(南京大学; 蚂蚁集团; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉语言模型不了解机器人身体运动关系的问题,提出KnowBody工具,通过可查询可修订的身体模型指导动作选择,在真实任务中显著提升完成率并减少规划轮次。
AI 中文摘要
通用视觉语言模型能够理解任务目标,却不知道特定机器人的运动及其功能部件如何产生预期效果。我们提出KnowBody,一种在保持模型权重冻结的同时,使这些与动作相关的身体关系变得明确、可查询且可修订的工具。从一个任务外轨迹初始化,部分身体模型指导动作选择并解释过去的交互。新证据会细化模型,且在重用前会重新检查依赖修订后身体估计的知识。在四个真实机器人任务的32次固定预算试验中,初始化的KnowBody实现了75%的完成率,而原生工具为25%,并且在两者均完成的任务中,成功试验所需的规划轮次更少。启用持续更新后,从第一次到第五次记录的成功,规划轮次减少了29%至53%。
英文摘要
A general-purpose vision-language model can understand a task goal without knowing how a particular robot's motion and functional parts produce the intended effect. We introduce KnowBody, a harness that makes these action-relevant body relations explicit, queryable, and revisable while keeping the model weights frozen. Initialized from one off-task trajectory, a partial body model guides action selection and the interpretation of past interactions. New evidence refines the model, and knowledge dependent on revised body estimates is rechecked before reuse. Across 32 fixed-budget trials on four real-robot tasks, initialized KnowBody achieves 75% completion versus 25% for the native harness and requires fewer planner rounds on successful trials in tasks completed by both. With persistent updates enabled, planner rounds decrease by 29-53% from the first to the fifth recorded success.
Comments19 pages, 5 figures, 7 tables. Project page: https://loule0-0.github.io/KnowBody/