读取房间:循环世界模型状态中的隐式困惑编码
Reading the Room: Implicit Confusion Encoding in Recurrent World Model States
查看机构详情
- University of Moratuwa(莫拉图瓦大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究发现RSSM架构世界模型(如DreamerV3)的隐藏状态$h_t$含隐式困惑信号,经线性探针、编辑验证其因果性,该信号可在多数控制任务中泛化。
中文摘要 AI 辅助
基于RSSM架构构建的世界模型(如DreamerV3)会保留一个循环隐藏状态$h_t$,该状态仅被训练以降低预测误差。我们发现该状态还会追踪自身的困惑,且这种困惑隐藏得极为隐蔽:它与$h_t$方差最大的方向几乎正交,任何基于方差的方法都无法检测到它。它在功能上与标记新输入的集成分歧、标记当前预测错误的重构误差截然不同。在保持预测误差固定而困惑度变化的测试中,对$h_t$的线性探针可检测到该信号(AUROC为0.72,共5次运行),而集成基准的得分低于随机水平。近期高误差步骤的 discounted count(折扣计数)可解释探针输出的80%($R^2=0.80$)。我们通过直接编辑$h_t$并观察行为变化,包括使用其他轨迹的真实值而非合成编辑进行验证,确认该信号是被因果使用的,而非仅存在于模型中。其几何结构和闭式形式可在三个控制任务间泛化;决定性的解离测试本身仅在一个任务上表现清晰,而其实际应用(决定何时核查现实而非依赖想象)可在三个任务中的两个任务上泛化。
英文摘要
World models built on the RSSM architecture, such as DreamerV3, keep a recurrent hidden state $h_t$ trained only to reduce prediction error. We show this state also tracks its own confusion, hiding in plain sight: nearly orthogonal to $h_t$'s directions of greatest variance, invisible to any variance-based method. It is functionally distinct from ensemble disagreement, which flags new inputs, and reconstruction error, which flags bad predictions right now. On a test holding prediction error fixed while confusion varies, a linear probe on $h_t$ finds the signal (AUROC 0.72, 5 runs), while an ensemble baseline scores below chance. A discounted count of recent high-error steps explains 80% of the probe's output ($R^2=0.80$). We confirm the signal is causally used, not merely present, by editing $h_t$ directly and watching behaviour change, including a check using real values from other trajectories instead of synthetic edits. Its geometry and closed form generalize across three control tasks; the decisive dissociation test itself holds cleanly on only one, and its practical use, deciding when to check reality instead of trusting imagination, generalizes to only two of the three tasks.