机器人学中的开环执行再探讨:迈向反应式、高性能策略
Revisiting Open-Loop Execution in Robotics: Toward Reactive, Higher-Performing Policies
浏览论文内容
中文总结 AI 辅助
本研究探讨机器人开环执行机制,发现长开环执行主要助短上下文策略模仿非马尔可夫示范,专家非马尔可夫性影响远超复合误差,长上下文下闭环策略更优,推动反应式策略范式发展。
中文摘要 AI 辅助
动作分块(即预测一系列动作并以开环方式执行前缀)已成为近期机器人操控模仿学习进展的关键推动因素。然而,执行过长的开环前缀会降低反应能力,限制策略修正误差的能力。此外,这些性能提升背后的机制仍知之甚少:现有研究将其归因于缓解复合误差、吸收推理延迟或平滑运动,但仅提供有限的可控证据,也未给出保持反应能力的指导。本研究认为,长开环执行主要是帮助短上下文策略模仿“非马尔可夫示范”。在4个仿真任务和2个真实世界任务中,我们表明专家非马尔可夫性强烈影响任务成功率与开环执行时长的关系。进一步,我们研究了复合误差(现有研究中长开环执行的主流解释)的影响,发现尽管复合误差有作用,但在我们的实验场景中,专家非马尔可夫性的影响大得多。最后,我们表明当策略具备足够长的上下文时,开环执行不再有益,最具反应性的闭环策略表现最佳。尽管模仿学习已通过长开环执行取得巨大成功,我们的发现推动长上下文、反应式策略成为更具原则性且性能更优的范式。
英文摘要
Action chunking --- the practice of predicting a sequence of actions and executing a prefix open-loop --- has emerged as a key enabler of recent progress in imitation learning for robotic manipulation. However, executing long open-loop prefixes reduces reactivity, limiting policies' ability to correct for errors. Further, the mechanisms underlying these performance benefits remain poorly understood: prior works cite mitigating compounding errors, absorbing inference latency, or smoothing motions, but provide limited controlled evidence or guidance for preserving reactivity. In this work, we argue that long open-loop execution primarily helps short-context policies imitate "non-Markovian demonstrations". Across four simulation and two real-world tasks, we show that expert non-Markovianity strongly shapes the relationship between task success and open-loop execution horizon. Further, we investigate the impact of compounding errors --- the prevailing explanation for long open-loop execution in prior work --- and find that while they matter, expert non-Markovianity has a much stronger impact in our experimental setting. Finally, we show that when policies are provided with a sufficiently long context, open-loop execution is no longer beneficial and the most reactive, closed-loop policies perform best. While imitation learning has seen great success using long open-loop execution, our findings motivate long-context, reactive policies as a more principled and performant paradigm.
发表机构
- Massachusetts Institute of Technology(麻省理工学院)
- Mundane Systems Inc.(芒德系统公司)
- University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。