视频记忆是否使用了其检索到的内容?记忆特异性的因果审计
Does Video Memory Use What It Retrieves? A Causal Audit of Memory Specificity
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Adobe Research(Adobe研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出读取时记忆替换方法,因果审计视频模型记忆特异性,发现记忆收益可源于通用表示、上下文或精确内容,为理解记忆机制提供直接区分手段。
AI中文摘要:
视频模型越来越多地使用记忆来在长序列中保留信息,其假设是性能提升来自于检索并利用正确的过去内容。标准的记忆消融实验测试记忆是否有帮助,但不测试检索到的内容是否真正起作用。我们通过读取时记忆替换直接测试这一点,该方法在保持计算其余部分不变的情况下替换被消费的记忆值。这区分了记忆收益与记忆特异性,即收益依赖于检索内容的程度。在冻结的视频世界模型中,不包含任何评估特定内容的身份无关控制,在Ego-Exo4D和7-Scenes上恢复了几乎全部收益,在TUM上恢复了约70%。在Ego-Exo4D的剂量反应中,随着这些值偏离观察到的训练记忆表示,恢复率从102%下降到1%,支持表示修复作为该设置中最受支持的解释。WorldMem显示出分级依赖性。来自同一轨迹的错误记忆相对于零内容恢复了94.1%的PSNR收益,而来自不相关轨迹和生物群的供体恢复了43.7%。SAM 2显示出强烈的内容依赖性。在DAVIS上,将正确的空间记忆替换为有效的错误记忆,使平均区域和边界分数从0.926降至0.182。在MOSEv2重现时,该分数从0.459降至0.000。这些结果表明,记忆收益可能依赖于通用表示支持、更广泛的上下文或精确的情景内容。读取时替换提供了一种直接区分它们的方法。
英文摘要:
Video models increasingly use memory to preserve information over long sequences, with the assumption that gains come from retrieving and using the correct past content. Standard memory ablations test whether memory helps, but not whether the retrieved content is responsible. We test this directly with read-time memory substitution, which replaces the consumed memory value while leaving the rest of the computation unchanged. This separates memory benefit from memory specificity, the extent to which the gain depends on retrieved content. Across frozen video world models, identity-free controls containing no evaluation-specific content recover essentially the full benefit on Ego-Exo4D and 7-Scenes and about 70% on TUM. In the Ego-Exo4D dose response, recovery falls from 102% to 1% as these values move away from observed training-memory representations, supporting representation repair as the best-supported explanation in this setting. WorldMem shows graded dependence. A wrong memory from the same trajectory recovers 94.1% of the PSNR benefit relative to zero content, while a donor from a disjoint trajectory and biome recovers 43.7%. SAM 2 shows strong content dependence. On DAVIS, replacing the correct spatial memory with a valid wrong memory reduces mean region and boundary score from 0.926 to 0.182. At MOSEv2 reappearance, it falls from 0.459 to 0.000. These results show that memory gains can depend on generic representation support, broader context, or exact episodic content. Read-time substitution provides a direct way to distinguish them.