arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RECAP-Forcing:保留内容外观以实现长视频生成

RECAP-Forcing: Retaining Content Appearances for Long Video Generation

Haiyang Xu, Zheng Ding, Zhuowen Tu

arXiv 2608.26671首次发表:更新:

发表机构

UC San Diego(加州大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长自回归视频生成的内存挑战,提出按外观新颖性组织内存的RECAP-Forcing方法,通过注意力汇和光流新颖性库统一机制,免训练且无额外参数,提升视频质量与语义保真度,优于现有内存方法。

AI 中文摘要

长自回归视频生成面临一个根本性的内存挑战:在有限的注意力窗口下,模型必须决定保留不断扩展的历史中的哪些信息。现有方法按时间组织内存,保留最近的帧,同时压缩或丢弃较旧的帧。我们反而提出RECAP-Forcing,按外观新颖性组织内存。长视频不仅仅是帧序列,还是一个不断演变的主体、对象和场景集合,其身份必须随时间保持一致。我们通过在新出现内容首次可见时,保留与新出现内容(如进入的主体、被遮挡后重新出现的区域和新引入的场景)相关的KV缓存来组织内存,优先考虑新颖性而非近期性。内存应随新引入内容的数量扩展,而非随视频长度扩展。这种按外观索引的内存使长程一致性成为内存结构的显式属性。我们的框架在这一单一原则下统一了两种机制:在视频开始时,当所有可见内容都是新颖的,一个注意力汇(attention sink)会保留初始场景;随着视频演变,一个基于光流的新颖性库(novelty bank)通过选择性保留新显现的内容来扩展相同原则。作为一种无额外可学习参数的免训练推理方法,RECAP-Forcing在多个强基准上持续提升视觉质量和语义保真度,且优于现有内存方法。

英文摘要

Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model must decide which information from an ever-expanding history to retain. Existing methods organize memory temporally, preserving recent frames while compressing or discarding older ones. We instead propose RECAP-Forcing, organizing memory by appearance novelty. A long video is not merely a sequence of frames, but an evolving cast of subjects, objects, and scenes whose identities must remain consistent over time. We organize memory by retaining the KV cache associated with newly appearing content--such as entering subjects, disoccluded regions, and newly introduced scenes--at the moment it first becomes visible, prioritizing novelty over recency. Memory should scale with the amount of newly introduced content, rather than with video length. This appearance-indexed memory makes long-range consistency an explicit property of the memory structure. Our framework unifies two mechanisms under this single principle. At the beginning of a video, when all visible content is novel, an attention sink preserves the initial scene. As the video evolves, an optical-flow-based novelty bank extends the same principle by selectively retaining newly revealed content. As a training-free inference method with no additional learnable parameters, RECAP-Forcing consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.

CommentsProject page: https://xxuhaiyang.github.io/RECAP-Forcing/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑