AI 中文总结
本文使用分层隐马尔可夫模型统一并扩展了基于目标的强化学习框架,解决了完全可观测性假设与外部目标提供的局限。
AI 中文摘要
Tasse等人(2026)提出的以智能体为中心的一般价值函数(ACGVF)构造方法,使得智能体能够做出通常由环境或智能体设计者强加的两个决策:追求哪个目标,以及何时宣布目标完成(此外还有选择动作)。这是一个非常通用的框架,几乎涵盖了先前关于强化学习、控制和规划的几乎所有工作,以及认知科学中提出的更一般的形式化方法。然而,它假设环境是完全可观测的,即观测是一个充分的统计量。在Murphy等人(2025)的研究中,提出了一种通用智能体设计,其中策略基于内部信念状态$z_t$和内部目标;然而,这些目标被假设为外部提供的。在本注记中,我们使用分层隐马尔可夫模型(HHMM)的形式化方法统一并扩展了这两种方法。
英文摘要
The agent-centric general value function (ACGVF) construction of \citet{tasse2026goal} lets the agent make two decisions that are normally imposed by the environment or agent designer: which goal to pursue and when to declare a goal as finished (in addition to choosing the action). This is a very general framework that subsumes almost all prior work on reinforcement learning, control and planning, as well as more general formalisms proposed in the cognitive sciences. However, it assumes the environment is fully observed, i.e., that the observation is a sufficient statistic. In \citet{murphy2025rl}, a general agent design was proposed where the policy is based on an internal belief state $z_t$ and an internal goal; however, the goals were assumed to be externally provided. In this note, we unify and extend these two approaches using the formalism of hierarchical hidden Markov models (HHMM) \citep{murphy2001hhmm}.