LayerRecall:用于视频生成中长时序一致性的状态条件化记忆路由
LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation
查看机构详情
- Zhejiang University(浙江大学)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出LayerRecall记忆路由与CHPM训练方法,在不牺牲局部连续性的前提下提升视频生成长程一致性,在MemoBench、MovieBench等数据集上取得最优结果,且具备跨主干可移植性与低推理开销。
中文摘要 AI 辅助
自回归视频扩散通过从有限的近期上下文生成片段,实现了可扩展的长视频生成。基于近因的缓存虽能保持局部连续性,但会在主体、对象、场景或属性再次出现时,驱逐所需的历史线索。现有记忆机制让模型接触到非局部历史,但仅访问并不能确保有效利用。我们的分析表明,视频DiT层对当前、近期和远程上下文表现出不同的偏好,这意味着长程记忆需要同时决定检索什么以及在哪里使用它。我们引入LayerRecall,一种当前条件化的、层选择性的记忆路由,它检索相关的历史K/V状态,并仅将其注入到特定于主干的记忆敏感层,同时在其他地方保留局部注意力。为减少对稀缺高质量长时序视频和显式内存分配标签的依赖,我们进一步提出了跨时序预测匹配(CHPM),它利用特权长上下文参考在预测空间中监督有限内存路由。在100个多镜头评估提示下,LayerRecall在MemoBench和MovieBench上取得了最佳整体结果,在VBench-Long上与其主干模型性能相当,证明其在不牺牲局部连续性的情况下具备更强的长程恢复能力。定性分析进一步揭示了记忆引导的自校正,即最初不匹配的局部属性会恢复到其历史外观,而无需重置正在进行的运动或场景结构。额外分析显示了跨主干的可移植性和可忽略的推理开销。
英文摘要
Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent context. While recency-based caching preserves local continuity, it evicts historical cues needed when subjects, objects, scenes, or attributes reappear. Existing memory mechanisms expose models to nonlocal history, but access alone does not ensure effective use. Our analysis reveals that video DiT layers exhibit distinct preferences for current, recent, and distant context, suggesting that long-range memory requires deciding both what to retrieve and where to use it. We introduce LayerRecall, a current-conditioned, layer-selective memory router that retrieves relevant historical K/V states and injects them only into backbone-specific memory-sensitive layers while preserving local attention elsewhere. To reduce reliance on scarce high-quality long-horizon videos and explicit memory-allocation labels, we further propose Cross-Horizon Prediction Matching (CHPM), which uses a privileged long-context reference to supervise the bounded-memory router in prediction space. Across 100 multi-shot evaluation prompts, LayerRecall achieves the best overall results on MemoBench and MovieBench while matching its backbone on VBench-Long, demonstrating stronger long-range recovery without sacrificing local continuity. Qualitative analyses further reveal memory-guided self-correction, whereby initially mismatched local attributes return to their historical appearance without resetting ongoing motion or scene structure. Additional analyses show cross-backbone portability and negligible inference overhead.