arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CrossEdit:跨模态训练实现丰富的音视频编辑

CrossEdit: Cross-Modal Training Enables Rich Audio-Visual Editing

William Chen, Prem Seetharaman, Ke Chen, Oriol Nieto, Kevin Duarte, Siddharth Srinivasan Iyer, Mamshad Nayeem Rizve, Zhiwen Cao, Shinji Watanabe, Yuanjun Xiong, Jianming Zhang, Zeyu Jin, Justin Salamon

arXiv 2610.10264首次发表:更新:

发表机构

Adobe Research; Carnegie Mellon University(Adobe研究院; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CrossEdit提出统一的全模态编辑模型,利用跨模态迁移实现零样本音视频电影场景编辑,通过合成数据微调和自监督重建学习指令遵循,并发布基准与指标验证效果。

AI 中文摘要

多媒体编辑要求生成模型能够解释复杂的组合式指令,并仅跨模态修改所需元素,这一能力在很大程度上超出了现有系统的范畴。当前方法通常局限于单一模态内的简单、单属性编辑,因为难以获得多样且高质量的编辑对,尤其是在跨模态任务中,对齐的音视频(AV)数据十分稀缺。我们提出了CrossEdit,一个统一的面向图像、音频和视频的全模态编辑模型,通过跨模态迁移实现零样本的AV电影场景编辑。我们的关键观察是,复杂的指令遵循能力,一旦在任一模态中习得,便能泛化到其他模态。我们利用声学信号的叠加性质,程序化地生成带有组合式指令的大规模音频编辑对,并表明在此合成数据上进行微调,结合自监督的AV掩码重建,以及一组精心策划的跨模态任务,能够诱导出可零样本迁移到未见模态和指令组合的指令遵循能力。为评估这一能力,我们发布了CrossEditBench,一个对电影场景进行AV编辑的人工标注基准,并提出了AV-FES,一个联合评分指令遵循与一致性的指标。我们表明,我们提出的技术提升了零样本AV电影场景编辑的性能,同时保持或提升了唇语同步语音编辑以及标准图像、视频和音频编辑基准的性能。演示可在该https URL获取。

英文摘要

Multimedia editing requires generative models to interpret complex, compositional instructions and modify only the desired elements across modalities, a capability largely beyond existing systems. Current approaches are typically limited to simple, single-attribute edits within one modality, since diverse, high-quality editing pairs are hard to obtain, especially for cross-modal tasks where aligned audiovisual (AV) data is scarce. We present CrossEdit, a unified omni-modal editing model for images, audio, and video that performs AV movie scene edits zero-shot through cross-modal transfer. Our key observation is that complex instruction-following, once learned in any modality, generalizes to others. We exploit the additive nature of acoustic signals to procedurally generate large-scale audio editing pairs with compositional instructions, and show that fine-tuning on this synthetic data and self-supervised AV masked reconstruction, alongside a targeted set of curated cross-modal tasks, induces instruction-following that transfers zero-shot to unseen modality and instruction combinations. To evaluate this capability, we release CrossEditBench, a human-annotated benchmark of AV edits on movie scenes, and propose AV-FES, a metric that jointly scores instruction following and consistency. We show that our proposed techniques improve performance on zero-shot AV movie scene editing, while maintaining or improving performance on lip-synced speech editing and standard image, video, and audio editing benchmarks. Demos are available at https://wanchichen.github.io/crossedit/.

Comments27 pages (10 main + 1 limitations); 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑