发表机构
Texas Tech University(德克萨斯理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出统一框架,将赋能(empowerment)目标与模型预测控制(MPC)结合,使用单一策略进行探测和行动,仅需一阶导数,在经典控制任务上验证了与任务成本结合可提升成功率。
AI 中文摘要
基于采样的模型预测控制(MPC)是轨迹优化的一种强大方法,但其性能依赖于信息丰富的成本函数,而这通常需要大量的领域知识来设计。另一种方法是直接从系统动力学中推导目标。赋能(Empowerment)定义为智能体输入与其未来状态之间的信道容量,提供了这样一种与目标无关的目标,并已被证明能在多个领域产生有用的行为。然而,现有的基于赋能(empowerment)的控制器有两个关键限制:赋能(empowerment)是使用一种与执行控制策略不同的虚拟探测策略来估计的,这模糊了其与最终行为之间的联系;并且其计算通常需要动力学的二阶导数,限制了与标准MPC方法的兼容性。我们使用单一策略进行探测和行动来表述赋能(empowerment),直接将赋能(empowerment)目标与执行行为联系起来。由此产生的目标仅需一阶导数,并且可以使用标准MPC求解器进行优化,既可以单独使用,也可以与任务特定成本结合使用。我们在经典控制任务上评估了所提出的方法,并表明将赋能(empowerment)与任务成本相结合,其成功率可达到或超过单独使用任一目标所达到的成功率。我们的公式为标准MPC与内在动机控制之间提供了实用的桥梁。
英文摘要
Sampling-based model predictive control (MPC) is a powerful approach to trajectory optimization, but its performance depends on an informative cost function that often requires substantial domain knowledge to design. An alternative is to derive objectives directly from the system dynamics. Empowerment, defined as the channel capacity between an agent's inputs and its future state, provides such a goal-agnostic objective and has been shown to produce useful behaviors across a range of domains. However, existing empowerment-based controllers have two key limitations: empowerment is estimated using a virtual probing policy distinct from the executed control policy, obscuring its connection to the resulting behavior; and its computation typically requires second-order derivatives of the dynamics, limiting compatibility with standard MPC methods. We formulate empowerment using a single policy for both probing and acting, directly linking the empowerment objective to the executed behavior. The resulting objective requires only first-order derivatives and can be optimized with standard MPC solvers, either alone or in combination with a task-specific cost. We evaluate the proposed approach on classical control tasks and show that combining empowerment with task costs matches or exceeds the success rate achieved by either objective alone. Our formulation provides a practical bridge between standard MPC and intrinsically motivated control.