发表机构
Max Planck Institute for Human Cognitive and Brain Sciences(马克斯·普朗克人类认知与脑科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究将主动推理构建为策略优化,证明其可表述为凸马尔可夫决策过程,分析了相关公式并推导MD算法,表明结合世界模型学习与策略优化能赋予主动推理执行性强化学习结构,为其奠定理论基础。
AI 中文摘要
主动推理(AIF)将适应性行为视为预期自由能(EFE)的最小化,在单一变分原理中结合认知和实用目标。我们将AIF构建为策略优化,并表明对于闭环控制策略,EFE最小化可被表述为凸马尔可夫决策过程(MDP)。在此表述中,实用项在预测状态边缘上是线性的,等同于潜在MDP中的奖励最大化,而认知值引入了非线性成分,将EFE最小化与标准强化学习区分开来。这一观点进一步揭示了主动推理的认知驱动作为与策略相关的(执行性)奖励。我们分析了EFE的有限时域、折扣和平均奖励公式,并推导了一种镜像下降(MD)算法,该算法围绕当前状态边缘对目标进行局部线性化,产生与演员-评论家方法和动态规划兼容的与策略相关的奖励。最后,我们认为将世界模型学习与策略优化相结合,赋予了主动推理执行性强化学习的结构,为在现代强化学习和优化理论中奠定主动推理基础提供了一条途径,包括收敛分析和有原则的策略改进保证。
英文摘要
Active Inference (AIF) frames adaptive behavior as the minimization of expected free energy (EFE), combining epistemic and pragmatic objectives within a single variational principle. We frame AIF as policy optimization and show that, for closed-loop control policies, EFE minimization can be formulated as a convex Markov decision process (MDP). This perspective reveals that policy-dependent reward prediction errors transmit natural gradients of the expected free energy backwards in time rather than up a hierarchy. Finally, we show that coupling world-model learning with policy optimization gives active inference the structure of performative reinforcement learning. Together this places EFE minimization within modern reinforcement learning and optimization theory and opens a route toward principled algorithms for active inference.