发表机构
King AI Labs, Microsoft Gaming; KTH Royal Institute of Technology(King AI实验室,微软游戏; 瑞典皇家理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出预测充分性指标,检验模型强化学习中上下文编码器的必要性,并给出实用标准以匹配上下文机制与可用信息。
AI 中文摘要
基于模型的强化学习中的泛化方法通常假设智能体无法从自身经验中恢复控制环境动态的潜在上下文,因此需要外部提供该上下文。我们通过“预测充分性”来形式化并检验这一假设,该指标量化了在智能体诱导的访问分布下,访问上下文对下一步预测的增量贡献,并将其分解为历史可恢复部分、需要真实上下文的残差部分以及有限模型带来的缺陷。我们根据条件集所能针对的预测风险对上下文感知算法进行分类,并在识别难度递增的环境中证明,这种余量并不遵循MDP类别。同一任务在不同先验下,一种设置中留下预测余量,另一种设置中则与零无显著区别,后者中智能体的行为隐式识别了上下文,且此类机制的任何益处不能归因于缺失信息。在余量持续存在的地方,学习到的状态仅部分暴露该余量,而添加真实上下文仍能降低风险。我们的贡献是一个实用标准,用于将上下文机制与其可用的信息匹配,该标准可从普通训练的智能体估计,无需参考策略。
英文摘要
Methods for generalization in model-based reinforcement learning typically assume that an agent cannot recover the latent context governing the environment dynamics from its own experience, and therefore supplies it externally. We formalize and test this assumption with \emph{predictive sufficiency}, which quantifies what access to the context adds to next-step prediction under the visitation distribution an agent induces, and separates that quantity into a history-recoverable part, a residual requiring the true context, and the deficit added by a finite model. We classify context-aware algorithms by the predictive risk their conditioning set can target and demonstrate across environments of increasing identification difficulty that the headroom does not follow the MDP class. The same task under different priors leaves predictive headroom in one setting and nothing distinguishable from zero in another, where the agent's behavior implicitly identifies the context and any benefit of such a mechanism cannot be attributed to missing information. Where headroom persists, the learned state exposes it only partially, and adding the true context still lowers the risk. Our contribution is a practical criterion for matching contextual mechanisms to the information available to them, estimated from the ordinary trained agent without a reference policy.