从1500小时演示数据及在线策略修正扩展双手机器人家务操作
Scaling Bimanual Household Manipulation from 1,500 hours of Demonstrations to On-Policy Corrections
浏览论文内容
中文总结 AI 辅助
本研究发布1500小时双手机器人家务操作演示数据,训练XR-2视觉-语言-动作模型,验证其学习能力与数据集的扩展潜力,开源数据支持相关研究。
中文摘要 AI 辅助
学习用于鲁棒双手机器人操作的通用策略受限于高质量大规模人类演示数据的稀缺性。本研究发布了涵盖日常家务任务的1500小时多样化双手操作演示数据,并利用该综合语料库训练XR-2模型,这是一种强大的视觉-语言-动作(VLA)模型。借助专门构建的高吞吐量数据管道和精心设计的多阶段训练范式,XR-2在系统实验中展现出优异的操作性能,同时保持了良好的训练效率和高数据利用率。我们进一步研究了两个关键扩展维度:不同规模的专家演示数据,以及基于实时人类干预的DAgger修正数据进行的后训练。在我们探索的所有数据范围内,两种设置下的任务成功率均稳步提升,在当前数据规模下呈现出清晰一致的扩展趋势。这些结果既验证了XR-2的学习能力,也证实了所发布数据集具有良好的扩展潜力,我们将其开源以支持双手机器人操作学习领域的可复现研究。
英文摘要
Learning generalist policies for robust bimanual manipulation is bottlenecked by the scarcity of high quality large scale human demonstration data. In this work, we release 1,500 hours of diverse bimanual manipulation demonstrations covering everyday household tasks, and use this comprehensive corpus to train XR-2, a powerful vision-language-action (VLA) model. Enabled by a purpose built high throughput data pipeline and a carefully designed multi stage training paradigm, XR-2 attains strong manipulation performance in our systematic experiments while retaining favorable training efficiency and high data utilization. We further study two critical scaling axes: varying the amount of expert demonstration data, and post training on DAgger correction data from real time human interventions. In both settings, task success rate improves steadily over the data ranges we probe, exhibiting a clear consistent scaling trend at our current data scale. These results validate both the learning capacity of XR-2 and the promising scaling properties of the released dataset, which we open source to support reproducible research on bimanual robot manipulation learning.
发表机构
- PrimeBot Research Institute(PrimeBot研究院)
- Swancor Advanced Materials Co., Ltd.(Swancor先进材料有限公司)
- School of Computer Science, Peking University(北京大学计算机学院)
- Crobotia
机构由 AI 辅助整理,请以论文原文为准。