arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编织强制:面向交互式长视频生成的组合记忆路由

Weave Forcing: Compositional Memory Routing for Interactive Long Video Generation

Ziyi Wang, Junchi Yao, Heqian Qiu, Wenbo Shi, Chengjiu Wang, Jinyang He, Binkai Hong, Hongliang Li

arXiv 2610.03510首次发表:更新:

发表机构

University of Electronic Science and Technology of China (UESTC); Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(电子科技大学; 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出编织强制框架,通过语义槽路由和掩码记忆编织实现组合记忆重用,提升交互式长视频生成的跨镜头一致性。

AI 中文摘要

自回归视频生成的最新进展改善了长时间跨度下的时间一致性,然而交互式叙事需要的不仅仅是连续场景的扩展:一个新镜头可能组合来自不同历史镜头中的角色和背景。整体提示检索可能忽略各组件不同的参考需求,而直接组合所有历史记忆则可能引入无关的视觉内容。为解决这些问题,我们提出了编织强制(Weave Forcing),一个无需训练的框架,用于交互式长视频生成中的组合记忆重用。首先,我们使用大语言模型(LLM)进行语义槽路由,将用户提示分解为角色和背景描述,并为每个组件明确选择合适的历史参考。为隔离所需内容,掩码记忆编织利用基于语义槽的条件对比注意力图来构建精细的语义掩码,有选择地暴露压缩历史键值(KV)记忆中的相关标记,以引导当前镜头的生成。我们进一步引入了覆盖自适应旋转位置编码(coverage adaptive RoPE),根据无、部分或完全参考覆盖来调整时间偏移和记忆保留,解决了当不完整的历史参考靠近当前生成位置时观察到的视觉伪影问题。大量实验表明,编织强制(Weave Forcing)在保持有竞争力的视觉质量和文本对齐的同时,改善了跨镜头主体和背景的一致性。

英文摘要

Recent advances in autoregressive video generation have improved temporal consistency over extended durations, yet interactive storytelling requires more than continuous scene extension: a new shot may combine characters and backgrounds from different historical shots. Whole prompt retrieval can overlook the distinct reference needs of individual components, while directly combining all historical memories may introduce unrelated visual content. To address these problems, we present Weave Forcing, a training-free framework for compositional memory reuse in interactive long video generation. First, we use an LLM for semantic slot routing to decompose user prompts into character and background descriptions and explicitly select suitable historical references for each component. To isolate the required content, masked memory weaving uses contrasting attention maps conditioned on semantic slots to construct refined semantic masks, selectively exposing relevant tokens from compressed historical KV memories to guide the generation of the current shot. We further introduce coverage adaptive RoPE to adjust temporal offsets and memory retention according to no, partial, or full reference coverage, addressing visual artifacts observed when incomplete historical references are positioned close to the current generation. Extensive experiments demonstrate that Weave Forcing improves cross-shot subject and background consistency while maintaining competitive visual quality and text alignment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑