arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28466cs.CV

过去塑造未来:自回归视频生成中的记忆

The Past Frames the Future: Memory for Autoregressive Video Generation

发表机构香港科技大学 · 香港城市大学 · 复旦大学
另 13 家 · 查看机构详情
  • HKUST(香港科技大学)
  • CityUHK(香港城市大学)
  • FDU(复旦大学)
  • ZODA
  • CMU(卡内基梅隆大学)
  • NYU(纽约大学)
  • HKUST(GZ)(香港科技大学(广州))
  • NUS(新加坡国立大学)
  • Georgia Tech(佐治亚理工学院)
  • PKU(北京大学)
  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)
  • NVIDIA(英伟达)
  • UCF(中佛罗里达大学)
  • UNITN(特伦托大学)
  • NTU(南洋理工大学)
  • UC Merced(加州大学默塞德分校)

机构由 AI 辅助整理,请以论文原文为准。

Harold Haodong Chen, Rongjin Guo, Disen Lan, Wen-Jie Shu, Hongfei Zhang, Hanzhe Hu, Shengtao Yao, Zixin Zhang, Guibin Zhang, Zhefan Rao, Jinxiu Liu, Yexin Liu, … 展开作者

Harold Haodong Chen, Rongjin Guo, Disen Lan, Wen-Jie Shu, Hongfei Zhang, Hanzhe Hu, Shengtao Yao, Zixin Zhang, Guibin Zhang, Zhefan Rao, Jinxiu Liu, Yexin Liu, Rui Peng, Yuhao Liu, Bin Ren, Shuai Yang, Yukang Chen, Salman Khan, Ying-Cong Chen, Ser-Nam Lim, Rynson W. H. Lau, Nicu Sebe, Yu Cheng, Ming-Hsuan Yang, Qifeng Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文系统综述自回归视频生成中的记忆机制,提出统一框架,从形式、功能、操作、学习与评估五视角分类,并指出可组合架构、可信更新等开放挑战。

中文摘要 AI 辅助

生成模型的进步提高了视频保真度,使得长时程生成、交互式世界建模和动态视觉环境成为可能。自回归(AR)视频生成通过因果推演扩展视觉序列。然而,一个根本性瓶颈随之出现:随着生成序列的扩展,实际模型必须在严格受限的上下文窗口、存储和计算限制下运行。因此,关键的历史信息,例如实体身份、动态状态和干预引起的因果变化,往往在其相关性减弱之前很久就离开了活动上下文。克服这一限制并维持时间持续性构成了一个基本的记忆问题。我们对AR视频生成中的记忆机制进行了系统而全面的综述。我们将记忆操作性地定义为跨外部AR步骤维持的持久历史信息,即使在原始证据不再局部可访问之后,它也能影响未来的生成。基于这一统一框架,我们通过五个互补视角组织文献:(I)形式,历史的表示载体;(II)功能,需要保留的特定语义和物理信息;(III)操作,记忆的写入、读取、更新、管理和整合的生命周期;(IV)学习,在闭环推演下优化记忆行为;(V)评估,诊断真正记忆能力的范式。最后,我们综合了开放性挑战,包括可组合和资源感知的记忆架构、可信的状态更新、自推演学习和标准化评估。通过连接表示、机制和学习范式,本文为开发可靠的、记忆条件下的视频生成系统奠定了结构化基础。

英文摘要

Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands, practical models must operate under strictly bounded context windows, storage, and computational limits. Consequently, critical historical information, e.g., entity identities, dynamic states, and intervention-induced causal changes, often leaves the active context long before its relevance diminishes. Overcoming this limitation and maintaining temporal persistence constitutes a fundamental memory problem. We present a systematic and comprehensive review of memory mechanisms in AR video generation. We formulate memory operationally as persistent historical information maintained across outer AR steps, capable of influencing future generation even after the originating evidence is no longer locally accessible. Building upon this unified framework, we organize the literature through five complementary perspectives: (I) Forms, the representational carriers of history; (II) Functions, the specific semantic and physical information requiring preservation; (III) Operations, the lifecycle of writing, reading, updating, managing, and integrating memory; (IV) Learning, the optimization of memory behaviors under closed-loop rollouts; and (V) Evaluation, the paradigms for diagnosing genuine memory capabilities. We conclude by synthesizing open challenges, including composable and resource-aware memory architectures, trustworthy state updating, self-rollout learning, and standardized evaluation. By bridging representations, mechanisms, and learning paradigms, this paper establishes a structured foundation for developing reliable, memory-conditioned video generation systems.

↑