发表机构
Adobe Research; Carnegie Mellon University(Adobe研究院; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CrossEdit提出统一的全模态编辑模型,利用跨模态迁移实现零样本音视频电影场景编辑,通过合成数据微调和自监督重建学习指令遵循,并发布基准与指标验证效果。
AI 中文摘要
多媒体编辑要求生成模型能够解释复杂的组合式指令,并仅跨模态修改所需元素,这一能力在很大程度上超出了现有系统的范畴。当前方法通常局限于单一模态内的简单、单属性编辑,因为难以获得多样且高质量的编辑对,尤其是在跨模态任务中,对齐的音视频(AV)数据十分稀缺。我们提出了CrossEdit,一个统一的面向图像、音频和视频的全模态编辑模型,通过跨模态迁移实现零样本的AV电影场景编辑。我们的关键观察是,复杂的指令遵循能力,一旦在任一模态中习得,便能泛化到其他模态。我们利用声学信号的叠加性质,程序化地生成带有组合式指令的大规模音频编辑对,并表明在此合成数据上进行微调,结合自监督的AV掩码重建,以及一组精心策划的跨模态任务,能够诱导出可零样本迁移到未见模态和指令组合的指令遵循能力。为评估这一能力,我们发布了CrossEditBench,一个对电影场景进行AV编辑的人工标注基准,并提出了AV-FES,一个联合评分指令遵循与一致性的指标。我们表明,我们提出的技术提升了零样本AV电影场景编辑的性能,同时保持或提升了唇语同步语音编辑以及标准图像、视频和音频编辑基准的性能。演示可在该https URL获取。
英文摘要
Multimedia editing requires generative models to interpret complex, compositional instructions and modify only the desired elements across modalities, a capability largely beyond existing systems. Current approaches are typically limited to simple, single-attribute edits within one modality, since diverse, high-quality editing pairs are hard to obtain, especially for cross-modal tasks where aligned audiovisual (AV) data is scarce. We present CrossEdit, a unified omni-modal editing model for images, audio, and video that performs AV movie scene edits zero-shot through cross-modal transfer. Our key observation is that complex instruction-following, once learned in any modality, generalizes to others. We exploit the additive nature of acoustic signals to procedurally generate large-scale audio editing pairs with compositional instructions, and show that fine-tuning on this synthetic data and self-supervised AV masked reconstruction, alongside a targeted set of curated cross-modal tasks, induces instruction-following that transfers zero-shot to unseen modality and instruction combinations. To evaluate this capability, we release CrossEditBench, a human-annotated benchmark of AV edits on movie scenes, and propose AV-FES, a metric that jointly scores instruction following and consistency. We show that our proposed techniques improve performance on zero-shot AV movie scene editing, while maintaining or improving performance on lip-synced speech editing and standard image, video, and audio editing benchmarks. Demos are available at https://wanchichen.github.io/crossedit/.
Comments27 pages (10 main + 1 limitations); 8 figures