世界模型更新的价值度量:面向持续适应的反事实效用协议
Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation
浏览论文内容
中文总结 AI 辅助
本文提出分叉账本协议,通过分支部署流比较更新与保持的反事实效用,证明固定更新机制降低控制任务回报,并支持以更新价值评估触发器。
中文摘要 AI 辅助
持续世界模型必须决定新数据是否值得改变模型。固定的回放调度和预测误差触发器指定了何时更新,但两者都无法揭示单次更新的价值:一次部署运行无法展示同一模型在保持其参数的情况下,在那一时刻本应表现如何。我们引入了分叉账本(fork ledger),该账本在预先登记的决策点将部署流分支为在公共随机数下的匹配更新和保持延续。它在相同的回合上评估两个延续,并记录 $\Delta R = R_{\mathrm{update}} - R_{\mathrm{hold}}$。始终应用一种固定的更新机制会降低所有三个模拟控制任务上的回报:CartPole($-144.0$;检查点自举 $95\\%$ 置信区间 $[-185.4,-116.1]$,相对于收敛回报约 $650$)、Walker($-82.8$;$[-101.1,-61.7]$)和 Cheetah($-18.6$;$[-29.0,-6.6]$)。发散是应用更新的结果,因此估计量计入每次尝试的分叉;限于 $720$ 次中未崩溃的 $693$ 次,CartPole 和 Walker 的符号不变($-113.4$ 和 $-82.1$),而 Cheetah 变得不明确($-3.9$;$[-17.5,+13.0]$)。任务是推断的单位:每个任务贡献 $240$ 次尝试的分叉,跨越五个预训练检查点与两个漂移方向交叉。账本使固定机制的反事实效用可观测,允许触发器根据其选择的更新来判断,而不仅仅是基于意外检测。
英文摘要
Continual world models must decide whether new data justify changing the model. Fixed replay schedules and prediction-error triggers specify when to update, but neither reveals the value of an individual update: one deployment run cannot show how the same model would have performed at that moment had it held its parameters. We introduce the fork ledger, which branches a deployment stream at pre-registered decision points into matched update and hold continuations under common random numbers. It evaluates both continuations on the same episodes and records $ΔR = R_{\mathrm{update}} - R_{\mathrm{hold}}$. Always applying one fixed update mechanism lowers return on all three simulated control tasks: CartPole ($-144.0$; checkpoint-bootstrap $95\%$ CI $[-185.4,-116.1]$, against a converged return near $650$), Walker ($-82.8$; $[-101.1,-61.7]$) and Cheetah ($-18.6$; $[-29.0,-6.6]$). Divergence is an outcome of applying the update, so the estimand counts every attempted fork; restricted to the $693$ of $720$ that did not collapse, CartPole and Walker are unchanged in sign ($-113.4$ and $-82.1$) and Cheetah becomes unresolved ($-3.9$; $[-17.5,+13.0]$). The task is the unit of inference: each contributes $240$ attempted forks over five pretrained checkpoints crossed with two drift directions. The ledger makes counterfactual utility observable for a fixed mechanism, allowing triggers to be judged by the updates they select rather than by surprise detection alone.
发表机构
- Substrate Labs
- UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。