arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01851cs.ROcs.AI

权重还是技能?机器人学习技术综述:从动作预测权重到自主编写技能的机器人

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

  • San Jose State University(圣何塞州立大学)
  • Meta
  • Apple(苹果公司)
  • Pragya Lab, BITS Pilani Goa(毕拉尼戈亚分校普拉gya实验室)

机构由 AI 辅助整理,请以论文原文为准。

Gaytri Jena, Kapil Wanaskar, Vinija Jain, Aman Chadha, Vasu Sharma, Amitava Das

AI总结:

本综述围绕机器人学习的“权重vs技能”轴,梳理两类技术的分类、自我提升机制与开放问题,考察77个代表性系统,为机器人学习领域提供分析框架。

AI中文摘要:

机器人学习正分化为两种方向:将能力嵌入冻结权重的策略(视觉-语言-动作模型,即VLA模型),以及能自行编写和优化可执行技能代码的智能体。本综述围绕“权重vs技能”这一轴组织该领域,核心分析贡献在于深入研究按自我提升程度排列的“代码即策略”方法:从零样本程序合成,到闭环自我修复与持久技能记忆,再到执行反馈、技能记忆与进化搜索结合为开放式循环的稀疏领域;仅有少数近期系统(如ASPIRE、ENPIRE和RoboClaw)处于该领域。我们梳理互补的“技能”端,从无监督强化学习技能发现到大语言模型技能库,发现“技能”一词至少有5种不同含义,其中仅代码类含义无需梯度更新即可自我提升。随后我们将该分类法与新兴的技能经济关联:商业机器人技能市场如今向机器人分发一键式技能,但仅提供静态回放,这凸显了适应性、跨实体可移植性、来源验证、安全性验证、组合性及标准化等开放问题。本综述聚焦性强,未详尽罗列领域,而是通过该分类法和一组对比表考察6个技术家族的77个代表性系统,提供了自我提升机制的操作定义,以及每个技术家族无法实现的功能说明。

英文摘要:

Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary "skills" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word "skill" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.

补充信息

↑