发表机构
Stanford University; Princeton University(斯坦福大学; 普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TANDEM结合任务与运动规划与按需人类演示,扩展规划域并高效收集演示,微调VLA模型,使五个长时程任务平均成功率从0%提升至60%。
AI 中文摘要
人类远程操作员花费大量时间演示机器人已经能够自主执行的行为,这限制了机器人基础模型数据收集的可扩展性。任务与运动规划(TAMP)可以自动化许多此类行为,但固定的规划域可能无法支持长时程操作任务的每个阶段。我们提出TANDEM(Tamp with As-Needed Demonstrations for Efficient Model fine-tuning),一个结合TAMP与选择性人类远程操作的系统,用于收集超出规划器能力范围的任务演示。我们的关键思想是将人类辅助表示为按需规划能力。给定语言指令和视觉观察,TANDEM使用预训练的视觉-语言模型扩展规划域,补充缺失的谓词和人类执行的魔法操作符。这使得规划器能够交错进行自主和人类执行阶段,而无需任务特定的干预点。在每个人类阶段之后,TANDEM重新感知场景,并在恢复自主规划前检查预期效果是否成立。为支持视觉-语言-动作(VLA)模型的微调,TANDEM还使用示例预训练轨迹,使规划器生成的运动与目标模型的预训练分布对齐。我们在五个超出TAMP域能力的长时程操作任务上评估TANDEM。在一个代表性的长时程任务上,在相同的人类干预时间内,TANDEM收集的演示数量是全任务远程操作的2.9倍。在五个任务上,每个任务使用20个TANDEM演示微调预训练的π_{0.5}-DROID模型,平均任务成功率从0%提升到60%。
英文摘要
Human teleoperators spend substantial time demonstrating behaviors that robots can already perform autonomously, limiting the scalability of data collection for robot foundation models. Task and motion planning (TAMP) can automate many of these behaviors, but a fixed planning domain may not support every stage of a long-horizon manipulation task. We present TANDEM (Tamp with As-Needed Demonstrations for Efficient Model fine-tuning), a system that combines TAMP with selective human teleoperation to collect demonstrations for tasks beyond the planner's capabilities. Our key idea is to represent human assistance as an on-demand planning capability. Given a language instruction and visual observation, TANDEM uses pretrained vision-language models to extend the planning domain with missing predicates and human-executed magic operators. This allows the planner to interleave autonomous and human-executed stages without task-specific intervention points. After each human stage, TANDEM re-perceives the scene and checks whether the intended effects hold before resuming autonomous planning. To support fine-tuning vision-language-action (VLA) models, TANDEM also uses example pretraining trajectories to align planner-generated motions with the target model's pretraining distribution. We evaluate TANDEM on five long-horizon manipulation tasks beyond the TAMP domain's capabilities. On a representative long-horizon task, TANDEM collects 2.9x as many demonstrations as full-task teleoperation at the same human intervention time. Fine-tuning a pretrained π_{0.5}-DROID model on 20 TANDEM demonstrations per task increases average task success from 0% to 60% across the five tasks.
CommentsUnder review. Project page: https://prpl-group.com/tandem/. The first two authors contributed equally