发表机构
Active Inference Institute(主动推理研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究主动推理智能体中无奖励预测组织与信息理论特征的关系,通过特定架构智能体在无奖励环境切换协议中测试,确定\(g\)为与\(\Phi_r\)相关时间组织的架构位置,反对将\(\Phi_r\)作为学习整合直接指标。
AI 中文摘要
最近一系列工作通过综合信息分解来衡量强化学习智能体中的因果涌现,报告称\(\Phi_r\)随训练增长并跟踪奖励改善。对于主动推理,这引发了无奖励预测组织如何与此类信息理论特征相关的问题。我在一个主动推理智能体中对此进行了测试,该智能体架构将快速感知潜在因素\(z\)与缓慢全局潜在因素\(g\)分离,其中\(g\)由预测误差驱动且与策略梯度在结构上解耦。在无奖励环境切换协议中,\(\Phi_r\)集中在\(g\)中;其总量在很大程度上取决于架构且随训练减少。学习的实质效果仅在原子组成层面清晰可见:解耦从负到正翻转符号并在环境变化下变得与环境无关,而向下因果关系携带依赖于环境的调整。这些结果将\(g\)确定为主动推理智能体中与\(\Phi_r\)相关的时间组织的架构位置,并反对将标量\(\Phi_r\)解读为学习整合的直接指标。
英文摘要
A recent line of work measures causal emergence in reinforcement learning agents through Integrated Information Decomposition, reporting that $Φ_r$ grows with training and tracks reward improvement. For active inference, this raises the question of how reward-free predictive organization relates to such information-theoretic signatures. I test this within an active inference agent whose architecture separates a fast perception latent $z$ from a slow global latent $g$, where $g$ is driven by prediction error and structurally decoupled from policy gradients. In a reward-free environmental regime-switching protocol, $Φ_r$ concentrates in $g$; its aggregate magnitude is largely architectural and decreases with training. The substantive effect of learning becomes legible only at the atom-compositional level: decoupling flips sign from negative to positive and becomes regime-invariant under environmental change, while downward causation carries the regime-dependent adjustment. These results identify $g$ as the architectural locus of $Φ_r$-relevant temporal organization in an active inference agent, and argue against reading scalar $Φ_r$ as a direct index of learned integration.