FineART:用于双臂操作任务的细粒度标注机器人轨迹数据集与视觉-语言-动作模型
FineART: Fine-Grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation
浏览论文内容
中文总结 AI 辅助
针对双臂长时程操作任务,提出密集标注数据集FineART及预测子任务的VLA模型FineART-VLA,中期训练显著提升成功率并降低数据需求。
中文摘要 AI 辅助
在真实环境中运行的机器人必须执行复杂、多步骤、长时间跨度的双臂任务,而非单一、孤立的动作。当前的操作数据集难以支持这种能力:尽管单臂数据集已达到数十万条轨迹的规模,但它们通常每条片段仅提供一个高层指令,而少数对子任务进行标注的双臂工作也仅标注了其时长的一小部分。我们提出了FineART,一个密集标注的双臂操作数据集,包含40,543条片段、1,718小时和533,913个子任务,覆盖151个任务。我们还引入了FineART-VLA,一种视觉-语言-动作策略,能够预测自身的下一个子任务,并展示了以这种方式进行中期训练能带来显著的性能提升。具体而言,在空间消歧任务上的成功率从32.0%提升至100.0%,而逐步的人类子任务引导将未见过的长时程任务的成功率从16.0%提升至76.0%。此外,在新机器人上进行最小微调后,该策略所需的数据量仅为未进行中期训练的基线方法的十分之一,并能零样本泛化到新硬件上完全未见过的任务。我们开源了完整的数据集、模型权重和训练代码。
英文摘要
Robots operating in real-world environments must often execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets provide limited support for this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode, while existing bimanual datasets provide subtask annotations for only part of their recorded hours. We present FineART, a densely annotated bimanual manipulation dataset comprising 40,543 episodes (1,718 hours) and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask to guide its actions. Mid-training on FineART's subtask annotations improves FineART-VLA's success at following spatial instructions from 32.0% to 100.0%. With step-by-step human subtask guidance, success on unseen long-horizon tasks increases from 16.0% to 76.0%. Furthermore, FineART-VLA matches baseline performance on a new robot with 10x less fine-tuning data and generalizes zero-shot to unseen tasks. We open-source the full dataset, model weights, and training code.
发表机构
- Scale AI
- Hugging Face
- University of Southern California(南加州大学)
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。