arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过无模型认知自由能估计器实现有原则的无方向内在动机

Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators

Alireza Furutanpey, Schahram Dustdar

arXiv 2607.16858首次发表:更新:

发表机构

Coovally; ICREA; Distributed Systems Group, TU Vienna(库瓦利; 加泰罗尼亚研究与高等研究院; 维也纳工业大学分布式系统组)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在含混合不确定性的环境中无监督强化学习的内在动机问题,提出基于无偏好预期自由能目标新颖性贡献的单一内在奖励,通过参数信息增益驱动探索,给出相关技术细节及成果,使智能体在不同状态区域表现良好。

AI 中文摘要

在具有混合不确定性来源的环境中,无监督强化学习需要不预先确定特定意外方向的内在动机。意外最小化设计用于‘不稳定’环境,预测误差好奇心奖励总预期意外,包括不可约噪声。强盗式或混合式在意外最小化和最大化奖励之间切换会重新引入非平稳性。我们提出一种在每个窗口内平稳的单一内在奖励,源自无偏好预期自由能目标的新颖性贡献。我们认为参数信息增益是状态空间高熵和低熵分量中合适的内在信号,最大化它能找到模型可解释的意外。在未解决动力学区域,认知项驱动探索;动力学解决时,认知项消失,同时偶然惩罚有利于低方差转换。伪计数提供认知价值,基于探针的惩罚捕获偶然方差,短视门保护信息丰富的后继状态。基于窗口冻结所有奖励定义对象可产生平稳的贝尔曼算子、学习目标的明确界限以及在混合、平滑度、带宽和容量假设下非参数估计器的条件均匀集中结果。从主动推理角度看,在保留新颖性时智能体无偏好,在完全可观测下标准似然模糊消失,添加非标准转换熵惩罚,且在状态空间已解决区域出现意外最小化。

英文摘要

Across environments with mixed sources of uncertainty, unsupervised reinforcement learning requires intrinsic motivation that does not precommit to a particular direction of surprise. Surprise minimization is scoped by design to ``unstable'' environments. Prediction-error curiosity rewards total expected surprise, including irreducible noise. Bandit or mixture switching between surprise-minimizing and surprise-maximizing rewards reintroduces non-stationarity by construction. We propose a single intrinsic reward, stationary within each window, derived from the novelty contribution of a preference-free Expected Free Energy objective, expressed in reward-maximization form. Our claim is that parameter information gain, the expected surprise of the next state minus its irreducible part, is the appropriate intrinsic signal in both high-entropy and low-entropy components of the state space. Maximizing it seeks exactly the surprise the model can explain away. In regions of unresolved dynamics, this epistemic term drives exploration. As dynamics become resolved, the epistemic term vanishes, while an aleatoric penalty favors lower-variance transitions, all without fitting an explicit next-state predictor. A pseudocount supplies epistemic value, a probe-based penalty captures aleatoric variance, and a short-horizon gate protects informative successors. A window-based freeze of all reward-defining objects yields a stationary Bellman operator, explicit bounds on learning targets, and a conditional uniform-concentration result for the nonparametric estimators under mixing, smoothness, bandwidth, and capacity assumptions. In active-inference terms, the agent is preference-free where novelty is retained, standard likelihood ambiguity vanishes under full observability, a nonstandard transition-entropy penalty is added, and surprise minimization emerges in resolved regions of the state space.

CommentsAccepted for a Spotlight Presentation at IWAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑