教学与成长:面向通用机器人学习的智能体中心架构
Teach and Grow: An Agent-Centered Architecture for General Robot Learning
查看机构详情
- School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究提出TGL智能体中心架构,通过将演示转化为可复用技能块,无需特定任务再训练即可实现通用机器人学习,在LIBERO评估中达SOTA性能,还提出相关缩放定律假设。
中文摘要 AI 辅助
端到端的视觉-语言-动作(VLA)和世界-动作模型为通用机器人技术提供了一条简洁路径,但它们的可靠性受限于已验证的物理覆盖范围。当陌生的物体、传感器、实体或接触超出该覆盖范围且不存在已验证的回退方案时,纠正故障需要新的机器人数据、策略更新和回归测试。这种反复出现的负担就是再训练成本。与文本不同,具身数据通常需要通过操作机器来生成。我们提出了教学与成长学习(Teach-and-Grow Learning,TGL),这是一种面向通用机器人学习的智能体中心架构。在通用形式中,多模态智能体将少量成功演示转化为可复用的技能块:针对有意义子目标的闭环行为。在新场景中,智能体会定位并组合这些技能块,选择已学习或几何工具,观察物理结果,并在执行偏离意图时修正路径。技能库存储可执行行为,而结构化经验记忆则记录成功、失败和修复情况。获取新任务无需针对特定任务进行策略再训练。我们在LIBERO评估中取得了最先进的性能;对照研究揭示了技能诱导、持久复用和智能体导向的自适应能力。最后,我们提出了教学与成长缩放定律假设:若X表示有效可复用经验,未来任务的误差和教学需求应作为X的幂律趋近于不可约下限。因此,该架构将部署视为持续学习的阶段,其中一项任务可让下一项任务更简单。
英文摘要
Vision-language-action (VLA) and world-action models typically absorb unfamiliar manipulation tasks through additional robot data collection and policy optimization. This recurring retraining burden slows the acquisition of new behavior. We present Teach-and-Grow Learning (TGL), a training-free architecture that turns a few successful demonstrations into reusable robot skills. Task acquisition requires no gradient updates, fine-tuning, or reinforcement learning: pretrained model weights remain fixed as the robot expands its explicit knowledge. Teaching is an accelerator, not a precondition, because the agent can also drive the robot directly, and demonstrations mainly improve reliability. Our implementation uses OpenAI GPT-6 Astra for multimodal reasoning and Codex to connect the agent to robot tools. The agent identifies subgoals shared across demonstrations, expresses them as closed-loop Skill Blocks, and grounds each block in the current scene. Physical feedback guides the next action and any recovery. Verified behaviors enter a persistent Skill Library; Experience Memory records the conditions and repairs that inform later decisions. TGL reaches 99.9% mean success on four LIBERO suites and 92.4% on seven LIBERO-Plus perturbation categories. Controlled studies show that taught blocks persist and improve related-task execution under the same model weights and executors. We further formulate a scaling hypothesis that relates effective reusable experience to falling future-task error and teaching demand. Code and demonstration videos: https://tgl.changnie.top