AI 中文总结
CAPEX利用基础模型作为自主演示者,通过经验自适应调整推理频率,高效生成机器人演示数据,使成功演示增加4.3倍、成本降低80%,训练的策略接近人类演示效果。
AI 中文摘要
机器人学习在很大程度上依赖于人类遥操作演示来获取有效的可学习行为。然而,人工操作的数据收集过程可能不直观、难以扩展,并且本质上是异步的。我们探索了一种替代方案:通过将基础模型本身用作自主演示者,从通用多模态基础模型中蒸馏物理行为到可部署的机器人策略。虽然足够强大的模型能够生成成功的零样本操作轨迹,但在物理执行过程中反复调用它们既缓慢又昂贵,限制了它们作为可扩展数据生成器的实用性。作为解决方案,我们引入了CAPEX,一种经验条件化的演示收集框架,它利用先前尝试的执行经验来调整基础模型必须观察、推理和重新规划的频率。我们在RoboCasa任务以及物理Franka和双臂YAM-arm平台上进行评估,测量任务成功率、模型调用次数、令牌使用量、收集时间和成本。我们进一步在匹配的人类遥操作和基础模型生成的演示集上训练Diffusion Policy和ACT,以评估自主收集数据的下游学习价值。我们发现,CAPEX将成功演示的数量增加了4.3倍,同时将每次成功演示的成本降低了80%。在CAPEX生成的数据上训练的策略接近在匹配的人类演示上训练的策略的性能;通过更长的训练,这种差距对于从头训练的策略来说基本消失。这些结果表明,基础模型可以作为可扩展的可重用机器人经验来源。项目页面:此https URL
英文摘要
Robot learning has largely relied on human-teleoperated demonstrations to acquire effective learnable behaviors. However, human-operated data collection processes can be unintuitive, difficult to scale, and inherently asynchronous. We explore an alternative: distilling physical behavior from general-purpose multimodal foundation models into deployable robot policies by using the foundation model itself as an autonomous demonstrator. While sufficiently capable models can generate successful zero-shot manipulation trajectories, repeatedly invoking them during physical execution is slow and expensive, limiting their utility as scalable data generators. As a solution, we introduce CAPEX, an experience-conditioned demonstration collection framework that uses execution experience from previous attempts to adapt how frequently the foundation model must observe, reason, and replan. We evaluate across RoboCasa tasks and on physical Franka and bimanual YAM-arm platforms, measuring task success, model calls, token usage, collection time, and cost. We further train Diffusion Policy and ACT on matched sets of human-teleoperated and foundation-model-generated demonstrations to evaluate the downstream learning value of autonomously collected data. We find that CAPEX increases the number of successful demonstrations by 4.3x while reducing the cost per successful demonstration by 80%. Policies trained on CAPEX-generated data approach the performance of those trained on matched human demonstrations; with longer training, this gap largely closes for policies trained from scratch. These results suggest that foundation models can serve as scalable sources of reusable robot experience. Project page: https://capex-paper.github.io/