发表机构
Stanford University; NVIDIA Research(斯坦福大学; 英伟达研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Robo-COP,通过部署中协同进化编排器与策略,解决VLA策略泛化瓶颈,提升模拟与真实任务的保留成功率。
AI 中文摘要
基于大型数据集训练的视觉-语言-动作(VLA)策略在其训练领域内表现出色,但仍无法泛化到机器人在现实部署中遇到的各种情况。智能体机器人系统通过视觉-语言模型(VLM)编排器来补充策略,该编排器学习何时调用策略、如何下达指令以及何时改用脚本化技能。然而,由于该框架围绕一个语言可操控性有限的冻结策略构建,编排器可以规避策略的失败,但永远无法克服这些失败。策略成为整个系统的瓶颈。微调策略可以消除这一瓶颈,但仅更新策略会使其与针对旧行为调优的编排器脱节。我们提出Robo-COP,其中编排器和策略在部署过程中协同进化。Robo-COP从自身执行中筛选技能演示,当这些数据能够解决重复出现的失败时微调策略,并且仅在新策略改进了其训练所针对的技能后才采纳它。在十个模拟RoboLab任务中,与使用冻结策略的相同框架相比,Robo-COP将平均保留成功率从64.8%提升至73.8%,而按固定计划微调且无验证的方法仅达到65.8%。在三个真实世界任务中,Robo-COP将保留成功率从38.3%提升至50.0%。Robo-COP将部署转变为自我改进的飞轮,机器人通过实践学习,每次执行中的改进都为下一轮学习产生更好的数据。视频和代码可在该https URL获取。
英文摘要
Vision-language-action (VLA) policies trained on large datasets are capable within their training domains, yet they still fail to generalize to the variety of situations a robot meets in real-world deployment. Agentic robot systems complement the policy with a vision-language model (VLM) orchestrator that learns when to call the policy, how to instruct it, and when to use scripted skills instead. However, because the harness is built around a frozen policy that has limited language steerability, the orchestrator can avoid the policy's failures but never overcome them. The policy becomes the bottleneck of the whole system. Fine-tuning the policy can remove this bottleneck, but updating it alone decouples it from an orchestrator tuned to its old behavior. We propose Robo-COP, in which the orchestrator and policy co-evolve during deployment. Robo-COP curates skill demonstrations from its own executions, fine-tunes the policy when this data can address recurring failures, and adopts each new policy only after it improves the skills it was trained for. Across ten simulated RoboLab tasks, Robo-COP raises mean held-out success from 64.8% to 73.8% over the same harness with a frozen policy, while fine-tuning on a fixed schedule without verification reaches only 65.8%. On three real-world tasks, Robo-COP raises held-out success from 38.3% to 50.0%. Robo-COP turns deployment into a self-improving flywheel in which robots learn by doing, with each improvement in execution producing better data for the next round of learning. Videos and code are available at https://robo-cop.pages.dev/.