arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于信念的最大占用原则与主动推断

Belief-Based Maximum Occupancy Principle and Active Inference

Manolis Mylonas, Rubén Moreno Bote

arXiv 2609.39342首次发表:更新:

发表机构

Center for Brain and Cognition; Universitat Pompeu Fabra(脑与认知中心; 庞培法布拉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将最大占用原则扩展至部分可观察环境,提出期望自由能的贝尔曼重构,通过信念状态值迭代实现离线计算,并在不确定食物源实验中比较发现MOP智能体在目标导向与探索间切换,而主动推断倾向单一食物源策略。

AI 中文摘要

内在动机在适应性和目标导向行为中扮演核心角色,通过赋予智能体与奖励无关的目标和对在嘈杂及不确定环境中行动有用的偏好,从而发挥作用。主动推断通过一个用于信念更新和行动选择的原理性框架,解决了在部分可观察环境中行动的问题。主动推断的一个关键组成部分是先前偏好的指定,它通过编码期望的未来结果来塑造行为。一种称为最大占用原则(MOP)的内在动机方法提出,智能体应行动以最大化未来状态和行动路径上的占用,而不带有任何偏好或认知目标。尽管其表述简单,MOP 却产生了丰富且适应性强的行为,这些行为将探索性变异与目标导向动态相结合。在这项工作中,我们将 MOP 扩展到部分可观察环境,并引入了主动推断的期望自由能的贝尔曼重构,两者都将基于信念的隐状态推断作为智能体状态的一部分。贝尔曼公式通过在整个信念状态空间上进行值迭代,实现了可处理的离线计算。我们在具有不确定食物来源的一组最小实验设置中比较了由此产生的行为。我们发现,MOP 智能体根据其可用能量和信念状态,在目标导向(寻找食物)行为与不同食物来源之间的探索之间切换。相比之下,主动推断智能体主要栖息在单一食物来源周围区域,这一策略具有较高的实用价值和认知价值。最后,我们与 Empowerment 进行比较,结果显示其与主动推断在性质上相似。

英文摘要

Intrinsic motivation plays a central role in adaptive and goal-directed behavior by conferring agents reward-independent objectives and biases useful to act in noisy and uncertain environments. Active Inference addresses the problem of acting in a partially observable environment through a principled framework for belief updating and action selection. A key component of Active Inference is the specification of prior preferences, which shapes behavior by encoding desirable future outcomes. An intrinsic motivation approach called the Maximum Occupancy Principle (MOP) proposes that agents act so as to maximize occupancy over future paths of states and actions, with no preferences or epistemic targets. Despite its simple formulation, MOP gives rise to rich and adaptive behaviors that combine exploratory variability with goal-directed dynamics. In this work, we extend MOP to partially observable environments and introduce a Bellman reformulation of the Expected Free Energy for Active Inference, both incorporating belief-based inference over hidden states as part of the agent state. The Bellman formulation enables tractable offline computation via value iteration over the full belief-state space. We compare the resulting behaviors in a set of minimal experimental settings with uncertain food sources. We find that MOP agents switch between goal-directed (food seeking) behavior and exploration between different food sources, depending on their energy available and their belief state. In contrast, Active Inference agents mostly inhabit regions around a single food source, a strategy having both high pragmatic and epistemic value. We finally compare with Empowerment, which is shown to be qualitatively similar to Active Inference.

CommentsAccepted at the 7th International Workshop on Active Inference (IWAI 2026, Madrid). To appear in Springer CCIS proceedings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑