被选择的未来仍可改写:视频模型中的因果可写性
A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models
浏览论文内容
中文总结 AI 辅助
本文发现视频模型生成错误运动并非未学到正确运动,而是未使用;通过低维物理变量编辑可恢复正确运动,即因果可写性,并在1.3B模型中验证其普遍性。
中文摘要 AI 辅助
当视频模型生成物理上不正确的运动时,是它未能学习到正确的运动,还是它学到了但未能使用?我们证明是后者:正确的运动仍然存在于模型内部,并且仍然可以被用来控制生成的视频。我们在红色质量缓慢振荡、蓝色质量快速振荡的视频上进行训练,然后测试一个具有快速观测运动的红色质量。即使模型在这种冲突情况下生成缓慢运动,一个从简单物理变量预测的低维编辑也能恢复正确的快速运动。我们将这种能力称为因果可写性。在固定强度下,我们发现一个清晰的深度边界:相同的编辑在边界之前改变视频,但在边界之后不改变。这种闭合标志着对该写入的承诺。然而,运动信号仍然存在,更强的下游写入可以恢复物理运动,而过度的增益会导致过冲。早期的因果可写性预测了训练后来纠正哪些错误:那些错误在更多网络深度上是可写的,而持续存在的错误则不是。我们在一个预训练的1.3B视频模型中复现了因果可写性及其清晰的闭合,支持了跨模型规模和训练机制的普遍性。
英文摘要
When a video model generates physically incorrect motion, did it fail to learn the correct motion, or did it learn it but fail to use it? We show the latter: the correct motion remains available inside the model and can still be made to control the generated video. We train on videos where red masses oscillate slowly and blue masses oscillate quickly, then test a red mass with fast observed motion. Even when the model generates slow motion in this conflicting case, a low-dimensional edit predicted from simple physical variables restores the correct fast motion. We call this ability causal writability. At fixed strength, we find a sharp depth boundary: the same edit changes the video before the boundary but not after it. This closure marks commitment for that write. The motion signal nevertheless remains, and a stronger downstream write can restore physical motion, while excessive gain overshoots. Early causal writability predicts which errors training later corrects: those errors are writable at more network depths than errors that persist. We reproduce both causal writability and its sharp closure in a pretrained 1.3B video model, supporting generality across model scale and training regime.
发表机构
- Tsinghua University(清华大学)
- Peking University(北京大学)
- University of Science and Technology of China(中国科学技术大学)
- MetaCircle
- Shanghai Qi Zhi Institute(上海期智研究院)
机构由 AI 辅助整理,请以论文原文为准。