发表机构
Universitat Pompeu Fabra(庞培法布拉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文重新审视稳态控制问题,提出用隐马尔可夫模型控制直接最大化期望观测概率,优于自由能变分方法,并给出最优策略确定性与最大占用原理的稳态新解。
AI 中文摘要
稳态的一种常见形式化是自由能原理,该框架定义了一组期望的观测值或临界状态,智能体应达到或保持接近这些状态。在自由能原理下,智能体应行动以最大化接收到期望观测的概率。在此,我们重新审视通过最大化变分下界(即所谓的负自由能)来求解最大化期望观测对数概率这一常见方法。我们表明,相反,直接在智能体策略下最大化该概率的方法更适合原始稳态控制问题,并为其提供更好的解决方案。这是通过隐马尔可夫模型(HMM)控制实现的,允许策略作用于隐藏状态或其噪声版本,同时试图最大化重复获得期望观测的概率。HMM控制大幅提升了相对于变分或自由能方法的性能。我们还表明,最优策略是严格确定性的,而变分方法导致随机策略近似。最后,我们提供了最大占用原理——该原理提出智能体应最大化动作-状态路径空间的占用——的稳态重新解释,通过将稳态定义为任何不会立即导致智能体终止或死亡的状态。
英文摘要
A common formalization of homeostasis is the free energy principle, a framework that defines a set of desired observation values, or critical states, that the agent should reach or remain close to. Under the free energy principle, an agent should act to maximize the probability of receiving the desired observations. Here we revisit the common approach of solving the problem of maximizing the log probability of the desired observations by maximizing a variational lower bound, the so-called negative free energy. We show that, instead, an approach directly maximizing that probability under the agent's policy is better suited to, and provides a better solution for, the original homeostatic control problem. This is done using hidden Markov model (HMM) control by allowing the policy to act over hidden states or noisy versions thereof while trying to maximize the probability of repeatedly having the desired observations. HMM control largely improves performance over the variational, or free energy, approach. We also show that the optimal policy is strictly deterministic, while the variational approach leads to a stochastic policy approximation. We finally provide a homeostatic reinterpretation of the maximum occupancy principle -a principle proposing that agents ought to maximize the occupancy of action-state path space -by defining homeostatic states as any states that do not immediately entail the termination or death of the agent.
Comments12 pages, 2 figures