发表机构
Sapienza University of Rome(罗马大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出将腿式机器人强化学习中的独立子策略组合作为独立研究问题,通过可靠技能组合实现可验证的机器人行为控制。
AI 中文摘要
机器人,尤其是人形机器人,在个体行为方面越来越有能力,每个行为都是通过训练专门的控制器获得的。专门技能训练速度快,收敛可靠,因为它面临的问题范围狭窄,并且可以独立验证,而单一端到端策略需要覆盖所有情况,这些优点它都不具备。仍然脆弱的是技能之间的转换。我们认为,独立子策略的组合本身就应该被视为一个研究问题,而不是作为实现细节留给任何碰巧可用的机制。可靠的组合是将一组独立技能转变为可被使用、扩展和共享的技能库的关键。更根本的是,如果控制可以在专门策略之间安全地传递,并且在任何时刻,机器人下一步应该做什么的选择都可以委托给一个性质完全不同的组件,如规划器、自动机或符号控制器,其行为可以事先检查。那么策略只需执行动作,而机器人可以被信任执行的任务将变得可验证。
英文摘要
Robots, and humanoid robots in particular, are increasingly competent at individual behaviors, each obtained by training a specialized controller. A specialized skill is quick to train, converges reliably because the problem it faces is narrow, and can be validated on its own, none of which is true of a single end-to-end policy asked to cover everything. What remains fragile is the transition between them. We argue that the composition of independent sub-policies deserves to be treated as a research problem in its own right, rather than as an implementation detail left to whatever mechanism happens to be at hand. Reliable composition is what turns a collection of separate skills into a repertoire that can be used, extended and shared. More fundamentally, if control can be passed between specialized policies safely, and at any moment, the choice of what the robot should do next can be delegated to a component of an entirely different nature, such as a planner, an automaton or a symbolic controller, whose behavior can be inspected in advance. The policies would then only ever have to act, and what the robot can be trusted to do would become verifiable.