什么阻碍了机器人递归自我改进?来自123轮智能体技能发现的经验教训
What Stops Recursive Self-Improvement in Robotics? Lessons from 123 Rounds of Agentic Skill Discovery
浏览论文内容
中文总结 AI 辅助
本研究通过123轮智能体技能发现实验,揭示机器人递归自我改进受阻于感知模块缺乏关系理解、技能链锁定早期步骤及评估框架误导,并提出相应改进建议。
中文摘要 AI 辅助
机器人能否像编码智能体改进软件那样自我改进?我们构建了一个智能体系统来探究这一问题。该系统观察机器人失败,找出缺失的能力,编写新技能或查找并安装外部模型,在仿真中测试每项更改,然后重复此过程,全程无需人类编写机器人代码。我们在家庭操作任务上运行了123轮改进。本报告描述了我们的发现。好消息是,该智能体能够自主发现能力:注意到目标超出视野后,它请求了一个主动查看模型,调试了该模型,并部署了一个可用的搜索技能。坏消息是,其改进并未累积成效。更改不断通过测试,但目标任务——将调味品放到冰箱顶层搁架上——从未成功。我们发现,智能体很少是瓶颈,其周围的三件事才是。第一,链式感知模块不理解关系。诸如SAM 3之类的分割器能识别搁架,但无法识别“顶层搁架”,因此智能体用越来越多的几何规则填补这一空白,却始终无法收敛,而它需要的是一种不同类型的模型。第二,技能链将学习锁定在第一步。长任务大多在早期失败,因此证据和修复都堆积在那里,后续技能很少被触及、测试或改进。第三,智能体所学内容由框架决定。智能体精确优化了评估者所衡量的内容,包括错误之处,而薄弱的测试和误导性的记忆使活动陷入停滞。我们将这些教训提炼为构建自我改进机器人系统的具体建议,每项建议均配有一个可能证伪它的实验。
英文摘要
Can a robot improve itself the way coding agents now improve software? We built an agentic system to find out. It watches a robot fail, works out which capability is missing, writes new skills or finds and installs external models, tests every change in simulation, and repeats, with no human writing robot code. We ran it for 123 improvement rounds on household manipulation tasks. This report describes what we learned. The good news is that the agent can discover capabilities on its own: noticing that its targets were out of view, it asked for an active-viewing model, debugged it, and deployed a working search skill. The bad news is that its improvements did not add up. Changes kept passing their tests, yet the target task, putting condiments on the top shelf of a fridge, never succeeded. We found that the agent was rarely the bottleneck. Three things around it were. First, chained perception modules do not understand relations. Segmenters such as SAM 3 find shelves but not "the top shelf", so the agent filled the gap with ever more geometric rules that never converged, when what it needed was a different kind of model. Second, skill chains lock learning onto the first step. Long tasks mostly fail early, so evidence and fixes pile up there, and later skills are rarely reached, tested, or improved. Third, what the agent learns is decided by the harness. The agent optimized exactly what the evaluator measured, including where it was wrong, and weak tests and misleading memory turned activity into a standstill. We distill these lessons into concrete recommendations for building robot systems that improve themselves, each paired with an experiment that could prove it wrong.