arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38839cs.CV

FrameMorrow:面向长时程视频生成的未来引导帧选择与前瞻令牌

FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation

Bo Yin, Xiaobin Hu, Jiaqi Zhao, Shuicheng Yan

首次发表
浏览论文内容

中文总结 AI 辅助

FrameMorrow提出基于未来信息需求的前瞻令牌选择历史帧,实现即插即用,在五个基准和11个模型上显著提升长程一致性、视觉质量与动作对齐。

中文摘要 AI 辅助

长时程视频生成要求模型有效利用不断增长的生成长度。随着生成历史的增长,保留所有先前内容变得日益昂贵且冗余,因此有效的历史选择至关重要。现有方法通常基于当前内容确定历史相关性。然而,与当前相关的信息不一定对未来生成有用,而看似不太相关的历史可能在未来变得重要。我们的关键见解是,应根据历史信息与未来信息需求的相关性来选择历史信息。捕捉这些需求并不需要生成完整的未来;相反,对未来重要内容的紧凑表示足以指导历史选择。基于这一见解,我们提出了FrameMorrow,一种前瞻性帧选择器,它预测一小部分代表未来信息需求的前瞻令牌,并利用它们从历史中识别相关信息。FrameMorrow选择显式的历史帧而非模型特定的内部状态,从而能够跨多种生成器实现即插即用集成,包括闭源模型,且额外推理成本极低。我们在五个基准和11个生成模型上评估了FrameMorrow,涵盖长视频生成、交互式生成和动作条件世界模型。大量实验表明,在各种生成设置中,长程一致性、视觉质量和动作对齐均得到持续改进。

英文摘要

Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retaining all previous content becomes increasingly expensive and redundant, making effective historical selection essential. Existing approaches often determine historical relevance based on the current content. However, information relevant to the present is not necessarily useful for future generation, while seemingly less relevant history may become important later. Our key insight is that historical information should be selected according to its relevance to future information needs. Capturing these needs does not require generating the full future; instead, a compact representation of what becomes important next is sufficient to guide historical selection. Building on this insight, we propose FrameMorrow, a prospective frame selector that predicts a small set of prospective tokens representing future information needs and uses them to identify relevant information from history. FrameMorrow selects explicit historical frames rather than model-specific internal states, enabling plug-and-play integration across diverse generators, including closed-source models, with little additional inference cost. We evaluate FrameMorrow across five benchmarks and 11 generative models spanning long-video generation, interactive generation, and action-conditioned world models. Extensive experiments demonstrate consistent improvements in long-range consistency, visual quality, and action alignment across diverse generation settings.

发表机构

  • National University of Singapore(新加坡国立大学)
  • Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

↑