arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18433cs.ROcs.LG

机器人基础模型中的具身差距

The Embodiment Gap in Robot Foundation Models

发表机构日本国立先进工业科学技术研究院(AIST)
查看机构详情
  • National Institute of Advanced Industrial Science and Technology (AIST)(日本国立先进工业科学技术研究院(AIST))

机构由 AI 辅助整理,请以论文原文为准。

Yukiyasu Domae, Keisuke Shirai, Hanbit Oh, Ryoichi Nakajo, Tomohiro Motoda, Koshi Makihara, Masaki Murooka, Takuma Yagi, Yoshiaki Bando, Ryo Hanai

首次发表
浏览论文内容

中文总结 AI 辅助

本综述定义机器人基础模型的具身差距,通过双轴图分类方法,从三个研究方向考察近期工作,提出报告框架以助力跨具身学习的评估与未来研究。

中文摘要 AI 辅助

机器人基础模型(Robot Foundation Models, RFMs),包括视觉-语言-动作(Vision-Language-Action, VLA)策略,常从缩放视角讨论:更多数据、更大模型、更广泛的基准应能提升泛化性。但在机器人领域,模型可能具备泛化性,却仍需大量工作才能在特定机身的机器人上运行。所需工作因方法和目标机器人而异,这些差异会影响实际部署。我们将可复用的模型、表示或数据与在目标机器人上执行应用之间的差距称为具身差距。本综述研究了机器人具身之间可复用的内容,以及新机器人上仍需实现的内容。我们将现有方法置于一个双轴图中,该图展示了共享结构的类型,以及在目标机器人上执行所需适应的阶段。随后,我们通过三个重叠的研究方向考察近期工作:共享语义与感知、共享机器人数据与接口、学习跨具身的对应关系。我们还提出了一种用于适应工作的报告框架,该框架仅靠成功率无法揭示,它能识别比较跨具身学习时应检查的工作,突出新机器人上仍需完成的工作,并指出未来研究的问题。

英文摘要

Robot foundation models (RFMs), including vision-language-action (VLA) policies, are often discussed through a scaling view: more data, larger models, and broader benchmarks should improve generalization. In robotics, however, a model can generalize while work still remains before it can run on a robot with a particular body. The work required differs across methods and target robots, and those differences affect practical deployment. We call the gap between reusable models, representations, or data and their use in execution on the target robot the embodiment gap. This survey examines what can be reused across robot embodiments and what must still be implemented on a new robot. We place existing methods on a two-axis map that shows the type of shared structure and the stage at which adaptation is needed for execution on the target robot. We then examine recent work through three overlapping research directions: sharing semantics and perception, sharing robot data and interfaces, and learning correspondence across embodiments. We also propose a reporting framework for adaptation work that success rate alone does not reveal. The framework identifies the work that should be checked when comparing cross-embodiment learning and highlights work that remains on a new robot and questions for future study.

补充信息

↑