作为rho-POMDPs中信念依赖效用的期望自由能
Expected Free Energy as Belief-Dependent Utility for rho-POMDPs
浏览论文内容
中文总结 AI 辅助
研究在部分可观测性下智能体信息收集决策问题,核心方法是用主动推理最小化期望自由能,此方法被证明等同于求解特定rho-POMDP,实验表明该方法在多种环境中表现良好,能提供基于信念的效用,避免过度探索。
中文摘要 AI 辅助
在部分可观测性下行动的智能体必须决定何时收集信息以及哪些观测值值得付出成本。标准POMDPs仅通过信息对奖励的最终影响来评估其价值。而rho-POMDP框架则通过依赖信念的效用rho直接奖励不确定性的降低,但在实践中,rho的选择及其权重都需要针对每个任务手动调整。我们表明主动推理完全消除了这种调整。最小化期望自由能(EFE)等同于求解一个效用为期望信息增益的rho-POMDP,并且探索权重固定为w = 1,因为变分界限以相同单位(奈特)表达了实用价值和认知价值。我们证明了对于观察然后决策的POMDPs的这种等价性,并将其扩展到因式观测POMDPs,这是一个更广泛的类别,涵盖了诸如无损检测和移动传感等交错的观察-行动问题,其中收集信息不会改变隐藏状态。实验支持了该理论。在从经典的老虎问题到RockSample以及一个具有超过65,000个状态的新结构检查基准等各种环境中,未调整的权重在相同的时间范围内匹配或优于仅基于奖励的规划,避免了针对每个任务调整奖励所导致的过度探索,并且位于成功-奖励帕累托前沿的奖励最大化拐点附近。实际的好处是一个现成可用的探索目标。在诸如故障检测和医疗筛查等应用中,每个测试都有成本,每个遗漏的故障都有代价,EFE提供了一种基于信念的效用,这种效用是推导出来的而不是调整出来的。
英文摘要
An agent acting under partial observability must decide when to gather information and which observations are worth their cost. Standard POMDPs value information only through its eventual effect on reward. The $ρ$-POMDP framework instead rewards uncertainty reduction directly, through a belief-dependent utility $ρ$, but in practice both the choice of $ρ$ and the weight placed on it are tuned by hand for every task. We show that active inference removes this tuning entirely. Minimizing Expected Free Energy (EFE) is exactly equivalent to solving a $ρ$-POMDP whose utility is expected information gain, and the exploration weight is fixed at $w=1$ because the variational bound expresses pragmatic and epistemic value in the same units (nats). We prove this equivalence for observe-then-commit POMDPs and extend it to factored observation POMDPs, a broader class that covers interleaved observe-act problems such as non-destructive testing and mobile sensing, where gathering information leaves the hidden state unchanged. Experiments support the theory. Across environments ranging from the classic Tiger problem to RockSample and a new Structural Inspection benchmark with over 65,000 states, the untuned weight matches or outperforms reward-only planning at the same horizon, avoids the over-exploration of bonuses tuned per task, and sits near the reward-maximizing knee of the success-reward Pareto frontier. The practical payoff is an exploration objective that works out of the box. In applications such as fault detection and medical screening, where every test has a price and every missed fault has a cost, EFE supplies a belief-dependent utility that is derived rather than tuned.
发表机构
- University of Colorado Boulder(科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。