发表机构
Mila – Quebec AI Institute; Université de Montréal(米拉-魁北克人工智能研究所; 蒙特利尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究任务从世界模型中所需的价值等价性,通过在DreamerV3堆栈上测量发现,潜在因素代表的闭包由训练目标维度决定,价值等价具维度性,单奖励目标是其一阶角,模型安装的任务结构与预测目标相关。
AI 中文摘要
一个经过学习的世界模型通常依据其对观测的忠实重建或奖励预测来评判,仿佛质量是模型简单具备或缺乏的东西。但任务实际从模型中所需的更窄:其查询所依赖的少数预测坐标,即我们所称的闭包。我们表明,一个潜在因素能代表多少闭包,并非由模型能力或观测决定,而是由其训练所针对目标的维度决定。我们在已知真实闭包的受控环境中,直接在DreamerV3堆栈上对此进行测量。一个对齐的标量值信号——价值等价核心的目标——仅安装了一个需要多个维度的闭包的一维投影:通过单个线性探针读取,当标量被完整目标取代时,可恢复结构从\(R^2 = 0.10\)升至\(0.76\)。将目标维度从一到四进行扫描,通过辅助头安装了恰好那么多预测方向,并且相同的阶梯状现象——幅度减弱但秩相同——通过模型自身的价值头也出现了,所以这种分离是维度上的而非头部形式的假象。能力匹配比较和现场压力检查排除了明显的替代方案。该规律支配着一个范围,我们测量了其边界:在一个结构可逐帧观测的伴随闭环任务中,重建安装了该结构且标量目标就足够了——在更便宜的训练信号无法恢复的地方,目标决定了潜在因素代表什么。因此,价值等价并非全有或全无,而是维度性的:熟悉的单奖励目标是其一阶角,并且模型安装的任务结构与要求其预测的目标一样多。
英文摘要
World models learn task-relevant information through many routes: observation reconstruction, recurrent state, temporal filtering, and explicit task supervision. Different routes can make different variables available. The same variable can also be available through several routes at once. When it is, looking at which route would increase the training loss most if removed does not tell you which route the model actually uses. The questions are reachability, whether a training signal can identify a task-relevant direction; admission, whether that direction is recoverable from the latent; and assignment, which eligible route carries it. We test them in environments with a known set of required coordinates. A direction cannot enter the latent unless some training signal can identify it. Reconstruction, recurrence, or filtering may already recover some of those coordinates; a reward or value head then has no residual direction to admit. For what remains, how many independent predictions the target supplies is how many coordinates install: one through four independent predictions admit one through four directions, including through the value head. Reachability is not admission: a temporal second-moment coefficient can remain absent under next-token prediction when accumulating it is a fraction of a percent of that loss, and a head that predicts the coefficient restores it. Assignment is a different test. Two routes that each carry the same variable when trained alone do not swap the carrier when we reverse which is more costly to remove. A recurrent model trained on a transformer's recorded sequences shows the same pattern. Near the point where the competing route is beginning to clear the probe threshold, independent training runs disagree. What a world model represents is therefore three questions: what information is reachable, what supervision admits, and which competing route carries it.
Comments24 pages, 12 figures. Substantial rewrite of v1 (title changed). v2 Adds the reachability / admission / assignment split and the assignment experiments / the route-boundary / RSSM-transfer / stochastic-onset and reachability-versus-admission experiments