发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SeeQ通过预测当前子任务的Q值来缩短信用分配跨度,利用离线数据中的子任务标注训练,并在测试时自回归预测子任务,从而提升长时程机器人操作策略的性能。
AI 中文摘要
尽管取得了快速进展,但通才机器人策略在复杂、长时程任务上仍然脆弱,这些任务包含多个阶段,或需要在同一基础阶段上反复尝试和深思熟虑后才能成功。Q值函数可以通过对候选动作进行排序或引导策略改进来提升这些策略,但从稀疏的任务级奖励中学习需要很长的信用分配跨度、困难的贝尔曼备份以及广泛的数据覆盖要求。我们提出了SeeQ(子任务引发的Q函数),它转而学习当前活动子任务的Q值。这缩短了价值预测的跨度,并使得使用时间差分(TD)目标进行有效学习成为可能。在训练期间,离线机器人数据中存在的子任务级标注提供了分解,并使得从广泛、可能次优的机器人数据集中学习成为可能。为了在测试时消除对人类标注或模块化子任务预测系统的需求,我们的Q函数架构被训练为在估计其价值之前,以自回归方式用自然语言预测活动子任务。我们使用基础视觉-语言骨干网络实例化SeeQ,在多样化的开源机器人操作数据上预训练,并在下游任务上进行微调。在两个双臂机器人平台上的四个真实世界操作任务中,SeeQ价值函数显著提升了最佳N选一策略引导的性能。
英文摘要
Despite rapid progress, generalist robot policies remain brittle on complex, long-horizon tasks that comprise multiple stages or require repeated attempts and deliberation on the same underlying stage before success. Q-value functions can improve these policies by ranking candidate actions or guiding policy improvement, but learning from sparse task-level rewards entails long credit-assignment horizons, difficult Bellman backups, and broad data-coverage requirements. We introduce SeeQ (Subtask-elicited Q-functions), which instead learns Q-values for the currently active subtask. This shortens the value-prediction horizon and enables effective learning with temporal-difference (TD) objectives. During training, subtask-level annotations present in offline robot data provide the decomposition and enable learning from broad, potentially suboptimal robot datasets. To eliminate the need for human annotations or modular subtask prediction systems at test time, our Q-function architecture is trained to autoregressively predict the active subtask in natural language before estimating its value. We instantiate SeeQ using a base vision-language backbone, pretrain it on diverse open-source robot manipulation data, and finetune it on downstream tasks. Across four real-world manipulation tasks on two bimanual robot platforms, the SeeQ value function substantially improves best-of-N policy steering.
CommentsWebsite: : https://saksham002.github.io/seeq/