arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DATAFARM: 面向视觉-语言-动作模型微调的分布对齐任务与运动规划

DATAFARM: Distribution-Aligned Task and Motion Planning for Fine-Tuning Vision-Language-Action Models

Samrat Sahoo, Yixuan Huang, Tom Silver

arXiv 2609.12316首次发表:更新:

发表机构

Stanford University; Princeton University(斯坦福大学; 普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对TAMP生成的轨迹与VLA预训练数据分布不匹配的问题,提出DATAFARM方法,在关节配置、运动风格和时间轮廓上对齐分布,显著提升微调效果,平均成功率从8.3%提升至56.7%,接近人类遥操作。

AI 中文摘要

收集高质量的机器人数据仍然是训练机器人基础模型的基本挑战。任务与运动规划(TAMP)提供了一种可扩展的演示生成方式,但我们的实验表明,尽管原始TAMP轨迹能够成功解决目标任务,但在用于微调预训练的视觉-语言-动作(VLA)模型时,其带来的益处却出奇地少。我们假设这一失败源于规划器生成的轨迹与用于预训练VLA的数据之间存在行为分布不匹配。为解决这一不匹配问题,我们提出了DATAFARM:用于微调机器人基础模型的分布对齐任务与运动规划,该方法将预训练分布直接纳入TAMP轨迹生成中。DATAFARM在机器人关节配置、运动风格和时间执行轮廓方面使生成的轨迹与预训练数据对齐。我们在三个TAMP可执行的桌面操作任务和一个超出TAMP能力的布料折叠任务上评估了DATAFARM。DATAFARM实现了56.7%的平均成功率,显著优于原始TAMP(8.3%),同时接近人类遥操作(61.7%)。在可变形物体操作任务上,该任务位于微调分布之外,微调后的模型保持了85%的成功率,而预训练模型为90%。这些结果表明,将规划器生成的演示与预训练分布对齐可以使TAMP成为VLA微调的有效数据来源。网站和代码:此https URL。

英文摘要

Collecting high-quality robot data remains a fundamental challenge for training robot foundation models. Task and motion planning (TAMP) offers a scalable way to generate demonstrations, but our experiments show that raw TAMP trajectories provide surprisingly little benefit when used to fine-tune pretrained vision-language-action (VLA) models, despite successfully solving the target tasks. We hypothesize that this failure arises from a behavioral distribution mismatch between planner-generated trajectories and the data used to pretrain the VLA. To address this mismatch, we introduce DATAFARM: Distribution-Aligned Task And motion planning for Fine-tuning A Robot foundation Model, an approach that incorporates the pretraining distribution directly into TAMP trajectory generation. DATAFARM aligns generated trajectories with the pretraining data in robot joint configurations, motion style, and temporal execution profiles. We evaluate DATAFARM on three tabletop manipulation tasks that TAMP can perform and a cloth-folding task beyond the capability of TAMP. DATAFARM achieves an average success rate of 56.7%, substantially outperforming raw TAMP (8.3%) while approaching human teleoperation (61.7%). On Deformable Object Manipulation, which is outside the fine-tuning distribution, the fine-tuned model retains 85% success, compared with 90% for the pretrained model. These results show that aligning planner-generated demonstrations with the pretraining distribution can make TAMP an effective source of data for VLA fine-tuning. Website and code: https://prpl-group.com/datafarm/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑