深度在组合中何时得以保留?潜在世界模型中的计算-质量机制
Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter
浏览论文内容
中文总结 AI 辅助
研究探讨潜在世界模型中深度在组合时的情况,用浅度惩罚ρ测试九个深度思维控制任务,发现三种机制,揭示反转机制由训练产生,其与观察/动作维度等相关,还指出任务机制受多种因素影响及相关注意事项。
中文摘要 AI 辅助
自适应计算世界模型,即早期退出或每步花费可变深度的深度混合预测器,假定深度能带来更好的预测且可自适应路由。在自回归展开中,第一个假设要求深度的每步精度在组合中得以保留。我们使用预先注册的工具——浅度惩罚ρ=误差(最浅退出展开)/误差(全深度展开),在九个深度思维控制任务上进行测试,涵盖匹配的单步(K=1)和多步(K=4)训练,每种情况有三个随机种子。我们发现三种机制:在6/9的任务中深度有助于展开(内在机制,ρ高达4.7倍),在2/9的任务中浅度退出优于全栈(反转机制,ρ低至0.85倍),还有一个是持平的。强大的反转机制(猎豹任务)并非动力学的属性,而是由训练产生:一种仅在首次展开步骤监督早期退出的消融操作会消除它(ρ从0.87变为1.18,n=8,Δ=+0.31),而一个内在权衡任务不受影响——我们将这种双重解离称为可路由性困境,因为使退出可路由的监督正是将它们训练得优于全栈的因素。该机制在一定程度上可先验预测:观察/动作维度和单步模型误差与ρ的斯皮尔曼相关性约为0.75(n=9)。在CEM规划器中,ρ的符号预测规划是否从深度中受益,在反转任务中最为明显,浅度规划优于深度规划。最后有三点注意事项:任务的机制取决于度量空间、展开视界和编码器。所有阈值和门在计算活动之前就已固定,包括对推动该研究的假设的预先注册的否定。
英文摘要
Adaptive compute for world models -- early-exit or mixture-of-depths predictors that spend variable depth per rollout step -- presumes that extra depth buys better predictions. In autoregressive rollouts, where planning actually happens, that premise requires depth's per-step precision to survive composition. We test it directly with one pre-registered instrument, the shallow penalty rho = err(shallowest-exit rollout)/err(full-depth rollout), on nine DeepMind Control tasks under matched single-step (K=1) and multi-step (K=4) training, eight seeds each. Three regimes emerge: depth helps (intrinsic, 6/9 tasks, rho up to 8x), depth actively hurts (inversion, 2/9, rho down to 0.87x), or depth barely matters (flat). The inversion is created by training, not the dynamics: supervising early exits only at the first rollout step erases it (Delta=+0.28, n=8, non-overlapping distributions) -- a routability catch-22: the per-step deep supervision that makes exits routable also trains them to out-roll the full stack. The regime is predictable: a frozen dimensionality-only classifier, committed before training, labels held-out tasks correctly out-of-sample, including an extreme extrapolation. The inversion reproduces under a transformer predictor, yet its manifestation is configuration-dependent, shifting with metric space, horizon, encoder, backbone, and -- most strongly -- training data: on the two tasks we retrained, competent-policy data removes both the inversion and the intrinsic tradeoff, loss unchanged. In a CEM planner, rho predicts whether planning benefits from depth. Every threshold and gate was committed before the corresponding compute, including a pre-registered negative for the motivating hypothesis. Whether more compute helps a world model is not a task property; it is a property of the operating configuration, with a stable, predictable, mechanism-backed core.
发表机构
- New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。