发表机构
The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出DREAM框架,可无需人类演示生成VLA微调数据,经实机实验验证其能降低新工作空间的数据收集成本并提升操纵任务成功率。
AI 中文摘要
视觉-语言-动作(VLA)模型在语言条件下的机器人操纵任务中已取得显著进展,但要提升其在新工作空间的性能,通常仍需要该环境中带动作标签的数据。通过人类遥操作收集此类数据成本高昂,尤其是当每个工作空间、物体摆放或任务可能需要新的演示时。我们提出DREAM,这一框架可从捕获的工作空间和语言指令中为预训练的VLA生成微调数据,无需特定任务的人类演示。DREAM会重建工作空间,利用大语言模型自动将指令转换为符号化的任务目标与成功标准,并运用任务与运动规划生成可行的机器人轨迹。规划得到的轨迹会在随机物体配置下进行扩充,经生成的成功标准验证后,渲染为用于VLA微调的图像-动作示例。我们通过语言条件操纵任务上的实机实验,探究DREAM能否作为部署工作空间的可扩展数据收集系统,具体考察在其自动生成的数据上进行微调是否比直接部署提升了成功率,以及将VLA适配到新工作空间时,其数据收集成本与人类遥操作相比如何。
英文摘要
Vision-language-action (VLA) models have made strong progress in language-conditioned robot manipulation, but improving their performance in a new workspace still often requires action-labeled data from that environment. Collecting such data by human teleoperation is costly, especially when each workspace, object arrangement, or task may require new demonstrations. We present DREAM, a framework that generates fine-tuning data for a pretrained VLA from a captured workspace and a language instruction, without requiring a task-specific human demonstration. DREAM reconstructs the workspace, automatically translates the instruction into symbolic task goals and success criteria using a large language model, and uses task-and-motion planning to generate feasible robot trajectories. The planned trajectories are augmented across randomized object configurations, verified by the generated success criteria, and rendered into image-action examples for VLA fine-tuning. Through real-robot experiments on language-conditioned manipulation tasks, we study whether DREAM can serve as a scalable data-collection system for the deployment workspace by examining whether fine-tuning on its automatically generated data improves success over direct deployment and how its data-collection cost compares with human teleoperation when adapting a VLA to a new workspace.
Comments8 pages, 5 figures