发表机构
University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对行为克隆在现实中的反直觉现象,提出OCBench基准,通过可控脚本化策略模拟人类示范特性,在受控环境中重现并科学探究这些现象。
AI 中文摘要
行为克隆(BC)尽管简单,但在现实世界中却表现出许多反直觉的现象。例如,随着模型对数据集过拟合程度的增加,BC的性能往往持续提升;而完全闭环的策略在没有动作分块的情况下往往会完全失败。遗憾的是,恰当地研究这些轶事性现象(“行为克隆之谜”)颇具挑战:在现实世界中,数据集和实验成本高昂且无法完全可控;在采用合成数据的仿真环境中,这些现象往往不易被观察到,部分原因是脚本化策略与人类示范之间存在差异。在这项工作中,我们提出了OCBench,一个具有可控脚本化策略的机器人操作基准,这些策略具有与人类示范相似的性质。我们表明,通过模仿人类示范的关键特性,OCBench在受控环境中重现了许多与BC相关的轶事现象。借助其GPU加速的环境和脚本化策略,我们展示了OCBench如何通过分析和驳斥各种假设,实现对先前报道的BC相关现象的科学探究。项目页面:此https URL
英文摘要
Behavioral cloning (BC), despite its simplicity, exhibits many counterintuitive phenomena in the real world. For example, the performance of BC often keeps increasing as the model overfits more to the dataset, and fully closed-loop policies often completely fail without action chunking. Unfortunately, properly studying these anecdotal phenomena ("behavioral cloning mysteries") is challenging: in the real world, datasets and experiments are costly and not fully controllable; in simulation with synthetic data, these phenomena are often not easily observed partly due to the discrepancy between scripted policies and human demonstrations. In this work, we propose OCBench, a robotic manipulation benchmark with controllable scripted policies that have similar properties to human demonstrations. We show that, by mimicking key properties of human demonstrations, OCBench reproduces many anecdotal BC-related phenomena in controlled settings. With its GPU-accelerated environments and scripted policies, we demonstrate how OCBench enables scientific studies of previously reported BC-related phenomena by analyzing and refuting various hypotheses. Project page: https://seohong.me/projects/ocbench