发表机构
Zhejiang University; Alibaba Group(浙江大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
InfinityEdit设计轻量型编辑适配器,解决无限视频编辑的延续性与稳定性问题,实现随编辑指令到来的视频流忠实编辑与无限生成。
AI 中文摘要
借助大型预训练模型,现有方法已有效改进了基于指令的视频编辑,但多数方法依赖原位编辑假设,在固定时间跨度内逐帧对齐编辑后的视频与给定源片段,该模式无法适用于开放式流场景,例如重新样式化直播游戏或对正在进行的镜头应用相机移动;此类场景中,编辑操作必须随着未来帧的到来而延伸,而非应用于静态输入片段。本文研究该设置并将其命名为无限视频编辑:给定前序片段与编辑请求,模型必须生成延续流的下一个片段,同时应用请求的编辑,随着无限序列的编辑指令到来,该过程会不断重复。此任务带来两项挑战:编辑必须是忠实的延续而非逐帧重写,且随着编辑累积,生成质量必须保持稳定。为应对这些挑战,我们首先设计了无限视频编辑的数据收集流程,基于收集的数据,提出InfinityEdit,这是一种轻量型编辑适配器,可为流式视频生成器赋予无限编辑能力。该适配器包含三个注意力模块:历史交叉注意力通过输入帧引导去噪帧,时间因果自注意力确保时间线索仅从较早帧流向较晚帧,编辑交叉注意力将编辑请求注入生成过程。推理期间,适配器仅在编辑请求到达的块中被激活,后续块由原始模型通过重置锚定帧生成,该方案在应用编辑的同时保留了原始模型的无限生成能力。大量实验表明,InfinityEdit可在每次编辑下忠实延续流,且在无限编辑序列中保持稳定。
英文摘要
With large pretrained models, existing methods have effectively improved instruction-based video editing. However, most of them rely on an in-place editing assumption. They align the edited video with the given source clip frame by frame over a fixed time span. This pattern fails for open-ended streams, e.g., restyling a live game or applying a camera move to an ongoing shot. In such cases, edits must extend to future frames as they arrive, rather than be applied to a static input clip. In this paper, we study this setting and name it infinite video editing: given a preceding segment and an edit request, a model must generate the next segment that continues the stream while applying the requested edit. This process repeats as an unbounded sequence of edit instructions arrives. This task brings two challenges: the edit must be a faithful continuation rather than a frame-wise rewrite, and generation quality must remain stable as edits accumulate. To address them, we first design a data-collection pipeline for infinite video editing. Based on the collected data, we propose InfinityEdit, a lightweight edit adapter that equips a streaming video generator with unbounded editing ability. The adapter contains three attention modules. History cross-attention guides the denoising frames using the input frames. Temporal causal self-attention keeps temporal cues flowing only from earlier frames to later ones. Edit cross-attention injects the edit request into generation. During inference, the adapter is activated only in the chunk where an edit request arrives. Subsequent chunks are generated by the original model with a reset anchor frame. This scheme applies the edit while preserving the original model's infinite generation ability. Extensive experiments show that InfinityEdit faithfully continues the stream under each edit, and stays stable over unbounded edit sequences.
Comments18 pages