联合与跨模态视频-音频生成及编辑:统一公式与设计分类法
Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy
浏览论文内容
中文总结 AI 辅助
本文提出统一公式与五轴分类法,系统梳理联合与跨模态视频-音频生成及编辑方法,首次将联合编辑划分为九类28种,并总结数据集、指标及关键开放问题。
中文摘要 AI 辅助
视频和音频通常被同时感知,然而大多数生成模型却孤立地处理它们。我们考察了联合建模这两种模态、从一种模态生成另一种模态,或以耦合方式编辑它们的方法,并围绕一个核心问题组织这些方法:输出如何在时间和语义上跨模态保持连贯?一个统一的公式将联合生成、跨模态生成和联合编辑定义为定义在音视频对单一分布上的三个问题,同时一个分类法沿五个设计轴比较了各类方法。据我们所知,这是首个系统性地对联合音视频编辑进行分类的综述,我们将其映射为涵盖28种编辑类型的九大编辑类别。我们针对每种设置描述了方法、数据集和指标,最后以我们认为最为重要的开放性问题作结。
英文摘要
Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics? A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems defined on a single distribution over audio-visual pairs, and a taxonomy compares methods along five design axes. To our knowledge, this is the first overview to systematically taxonomize joint audio-visual editing, which we map as nine edit categories spanning 28 edit types. We describe methods, datasets, and metrics for each setting and close with the open problems we view as most consequential.
发表机构
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
- University of Texas at Dallas(德克萨斯大学达拉斯分校)
- Virginia Tech(弗吉尼亚理工大学)
- Texas A&M University(德克萨斯A&M大学)
- University of Southern California(南加州大学)
- Vanderbilt University(范德堡大学)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Adobe Research(Adobe研究院)
- Arizona State University(亚利桑那州立大学)
- University of Oregon(俄勒冈大学)
- Stanford University(斯坦福大学)
- Dolby Laboratories(杜比实验室)
- Cisco(思科)
机构由 AI 辅助整理,请以论文原文为准。