FlowMimic:基于像素对扭曲流场的无掩码视觉编辑与生成,用于在线视频编辑数据生成及模态模仿
FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry
浏览论文内容
中文总结 AI 辅助
研究探索在单一模型中整合视频和图像模态的生成与编辑能力,开发像素对时间扭曲流场实时生成视频编辑样本,设计模态模仿损失对齐模态能力,引入相关任务及损失让模型内化基于语言的视觉编辑能力。
中文摘要 AI 辅助
随着视觉研究的主流趋势,我们探索在单一模型中整合视频和图像模态的生成与编辑能力。当前收集视频编辑数据的方法依赖劳动密集、耗时的策划过程,任务扩展性有限。我们开发了像素对时间扭曲流场,能从图像编辑样本实时生成视频编辑样本,并证明模型仅用此类数据就能学习视频编辑。我们将图像模态视为视频模态的特殊形式,设计模态模仿生成损失和模态模仿编辑损失来对齐两种模态的能力。此外,基于语言的视觉编辑需要理解编辑指令和参考视觉内容等,现有方法多依赖外部辅助,我们引入相关任务及损失让模型内化此能力。
英文摘要
In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting video editing data typically depend on labour-intensive, time-consuming curated procedures--involving object mask annotation, the use of error-introducing pair synthesis via I2V model and ControlNet-like guidance, and VLM-based quality filtering or refinement--and demonstrate limited task scalability. As a result, the diversity of editing tasks remains substantially narrower than that available for image editing models. We develop a pixel-pair temporal warped flow field that can directly generate corresponding video editing samples in real time from image editing samples, and we demonstrate across multiple levels of video editing tasks that a model can learn video editing using only such data. We regard the image modality as a particular form of the video modality. Accordingly, we design a modality mimic generation loss and a modality mimic editing loss to relatively align the capabilities--and thereby the output distributions--of the two modalities through mutual imitation. Moreover, language-based visual editing entails the comprehension of the editing instruction and the reference visual content, the localization of the region corresponding to that instruction within the reference visual contents, and the modification of that region alone. Existing approaches predominantly rely on external aids, such as fine-tuning an additional MLLM or explicitly supplying a mask sequence as auxiliary input during inference. In contrast, we aspire for the model to internalize this capability. To that end, we introduce sense-related tasks--for instance, referring expression segmentation--along with corresponding editing-region-aware latent-level loss and attention-level loss.