arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36416cs.ROcs.AIcs.CVcs.LG

FineART:用于双臂操作任务的细粒度标注机器人轨迹数据集与视觉-语言-动作模型

FineART: Fine-Grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation

Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci, Aseem Doriwala, Catherine Weaver, Mouli Sivapurapu, Kai Yang, Thomas Wolf, Jackson Lee, Pragna Mannam

首次发表
浏览论文内容

中文总结 AI 辅助

针对双臂长时程操作任务,提出密集标注数据集FineART及预测子任务的VLA模型FineART-VLA,中期训练显著提升成功率并降低数据需求。

中文摘要 AI 辅助

在真实环境中运行的机器人必须执行复杂、多步骤、长时间跨度的双臂任务,而非单一、孤立的动作。当前的操作数据集难以支持这种能力:尽管单臂数据集已达到数十万条轨迹的规模,但它们通常每条片段仅提供一个高层指令,而少数对子任务进行标注的双臂工作也仅标注了其时长的一小部分。我们提出了FineART,一个密集标注的双臂操作数据集,包含40,543条片段、1,718小时和533,913个子任务,覆盖151个任务。我们还引入了FineART-VLA,一种视觉-语言-动作策略,能够预测自身的下一个子任务,并展示了以这种方式进行中期训练能带来显著的性能提升。具体而言,在空间消歧任务上的成功率从32.0%提升至100.0%,而逐步的人类子任务引导将未见过的长时程任务的成功率从16.0%提升至76.0%。此外,在新机器人上进行最小微调后,该策略所需的数据量仅为未进行中期训练的基线方法的十分之一,并能零样本泛化到新硬件上完全未见过的任务。我们开源了完整的数据集、模型权重和训练代码。

英文摘要

Robots operating in real-world environments must often execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets provide limited support for this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode, while existing bimanual datasets provide subtask annotations for only part of their recorded hours. We present FineART, a densely annotated bimanual manipulation dataset comprising 40,543 episodes (1,718 hours) and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask to guide its actions. Mid-training on FineART's subtask annotations improves FineART-VLA's success at following spatial instructions from 32.0% to 100.0%. With step-by-step human subtask guidance, success on unseen long-horizon tasks increases from 16.0% to 76.0%. Furthermore, FineART-VLA matches baseline performance on a new robot with 10x less fine-tuning data and generalizes zero-shot to unseen tasks. We open-source the full dataset, model weights, and training code.

发表机构

  • Scale AI
  • Hugging Face
  • University of Southern California(南加州大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑