AI 中文总结
该研究揭示动作分块提升机器人控制性能的原因是隐式集成,提出带随机延迟的策略集成可匹配其性能,显式集成策略类在多领域表现更优。
AI 中文摘要
动作分块——即预测并执行多个动作而非单个动作——已被证明是学习有效机器人控制策略的关键组成部分。然而,我们对动作分块提升性能的原因仍缺乏精准理解。本研究旨在填补这一空白。通过在模拟和真实场景下的严格实验评估,我们发现关于动作分块成功原因的现有假设——时间一致性、 horizon 缩减和表示学习——均无法解释其成功。相反,我们发现与马尔可夫策略相比,动作分块得益于更强的非马尔可夫表达能力和更小的复合误差,但在许多关注场景中,这些效果可完全由延迟策略捕捉,延迟策略每一步基于过去k步的观测预测单个动作。我们进一步发现动作分块存在额外益处,称之为隐式集成。具体而言,通过学习多种时间关系(即a_t | o_t, a_t | o_{t-1}等),动作分块策略表现出与模型集成相匹配的行为,相比仅学习单一时间关系的策略,其鲁棒性和泛化能力更强。基于这些见解,我们证明在模拟和真实机器人控制场景中,可通过将动作分块策略部署为带随机延迟的策略集成,在不使用动作分块的情况下达到其性能。此外,我们提出一种策略类,通过显式实例化集成来放大动作分块的益处,且在多个领域中其性能显著优于动作分块。
英文摘要
Action chunking---predicting and executing multiple actions instead of a single action---has proven to be a critical component for learning effective robotic control policies. However, our precise understanding of why action chunking improves performance has remained limited. In this work we seek to close this gap. Through rigorous experimental evaluations in both simulated and real-world settings, we show that existing hypotheses for the success of action chunking---temporal consistency, horizon reduction, and representation learning---fail to explain the success of action chunking. Instead, we find that action chunking benefits from greater non-Markovian expressivity and reduced compounding error compared to Markovian policies, but, in many settings of interest, these effects can be fully captured by delayed policies, which at each step predict a single action based on the observation $k$ steps in the past. We then show that there exists an additional benefit of action chunking that we refer to as implicit ensembling. In particular, by learning a diversity of temporal relationships (that is, $a_t | o_t, a_t | o_{t-1}, \ldots$), action-chunked policies exhibit behavior matching that of a model ensemble, increasing their robustness and generalization ability over policies that only learn a single temporal relationship. Building on these insights, we show that in simulated and real-world robotic control settings, we can match the performance of action chunking without action chunking---by deploying an action chunking policy as an ensemble of policies with randomized delays. Furthermore, we propose a policy class that amplifies the benefits of action chunking by explicitly instantiating an ensemble, and which we show significantly improves over the performance of action chunking in many domains.