发表机构
MBZUAI; UESTC; Dalian University of Technology(穆罕默德·本·扎耶德人工智能大学; 电子科技大学; 大连理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RolloutFaith框架,审计视觉世界模型中内部干预的持久性,发现现有编辑器长期效果有限,并验证训练应奖励未来后果的假设。
AI 中文摘要
探针、激活补丁和学习型编辑器等可解释性方法旨在揭示或修改模型的当前计算。世界模型提出了更严格的要求:因为其预测会成为后续预测的输入,所以有用的内部修正必须在编辑停止后仍然有效。因此,我们提出了RolloutFaith,一个在干预时产生的预测以及后续在固定事件、动作、噪声和信息预算下的自主预测中衡量语义改进的框架。我们在Crafter、Cartpole和CoinRun上对三个世界模型评估了十个拟合编辑器。我们还使用了参考激活补丁(Reference Activation Patching),该方法将模型激活替换为从真实观察中计算出的配对激活,以衡量在所选接口处可用的修正。这种参考干预在全部九个模型和任务组合中改善了后续预测,并在八个组合中优于最佳拟合编辑器,然而其持续增益在九个组合中的五个里随视界增加而下降。当前拟合编辑器仅能恢复有限且不一致的长期效果。通过将各个状态组件恢复到其未触及的值,我们发现持久效应在DIAMOND中通过最新生成的帧传播,在DreamerV3中通过循环记忆传播,在STORM中则通过两者传播。这些发现表明训练应奖励未来的后果。为验证这一假设,我们提出了延迟LoReFT(Delayed LoReFT),它通过四个冻结的未来转移优化相同的低秩干预,并在一定程度上改善了持续干预效果。
英文摘要
Interpretability methods such as probes, activation patches and learned editors are designed to reveal or modify a model's current computation. World models pose a harder requirement: because their predictions become inputs to later predictions, a useful internal correction must survive after editing stops. We therefore propose RolloutFaith, a framework that measures semantic improvement both in the prediction produced at intervention time and over later autonomous predictions under fixed events, actions, noise, and information budgets. We evaluate ten fitted editors on three world models across Crafter, Cartpole, and CoinRun. We also use Reference Activation Patching, which replaces a model activation with the paired activation computed from the real observation, to measure the correction available at the chosen interface. This reference intervention improves later predictions in all nine model and task combinations and outperforms the best fitted editor in eight, yet its sustained gain decreases with horizon in five of nine combinations. Current fitted editors recover only limited and inconsistent long term effects. By restoring individual state components to their untouched values, we find that persistent effects travel through the newest generated frame in DIAMOND, recurrent memory in DreamerV3, and both in STORM. These findings suggest that training should reward future consequences. To test this hypothesis, we propose Delayed LoReFT, which optimizes the same low rank intervention through four frozen future transitions and improves sustained intervention effects to some extent.