SepGen:单模型中的多声源音视频分离与生成
SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model
浏览论文内容
中文总结 AI 辅助
SepGen 是一款单模型,可在一次采样中输出视频、混合 soundtrack 及各声源波形,支持生成与分离模式,在语音分离等任务上优于语言条件分离器。
中文摘要 AI 辅助
4D音视频场景包含一段视频、其描绘的动态几何结构以及场景中的声源,每个声源都有各自的位置和轨迹。从新视角渲染此类场景需要每个声源都以单独的波形形式存在,以便在信号混合前将其定位在场景中并传播到观察者。联合音视频生成器可以合成视频及其 soundtrack,但 soundtrack 以单一音频混合形式输出,无法单独访问各个声源。我们提出 SepGen,它扩展了预训练的音视频生成器,可在一次联合采样运行中输出视频、混合 soundtrack 以及每个带字幕声源对应的一个波形。SepGen 支持两种互补模式:生成模式和分离模式。在生成模式下,每个声源字幕指定其 stem 包含的内容,例如两人对话会按原轮次顺序输出每个说话人对应的一个 stem。在分离模式下,输入音频混合保持干净,而字幕指定要提取的内容,因此模型可根据自由文本描述分解录音。我们在文本合成的场景上评估生成任务,在其他生成器渲染的场景以及语音、音乐和音效的真实录音上评估分离任务。给定音频混合和包含语音台词的字幕,SepGen 的性能优于语言条件分离器,在语音任务上表现最为明显,且当从字幕中移除台词时仍保持领先。代码、检查点和数据集可在该 https URL 获取。
英文摘要
A 4D audio-visual scene comprises a video, the dynamic geometry it depicts, and the sound sources that populate it, each with its own position and trajectory. Rendering such a scene from a novel viewpoint requires that every source be available as an individual waveform, so that it can be localized in the scene and propagated to the observer before the signals are mixed. Joint audio-video generators can synthesize the video and its soundtrack, but the soundtrack is emitted as a single audio-mix in which the sources are not individually accessible. We present SepGen, which extends a pretrained audio-video generator to emit the video, the mixed soundtrack, and one waveform per captioned source in one joint sampling run. SepGen supports two complementary modes: generation and separation. In generation mode, each source caption specifies what its stem contains. A two-speaker dialogue, for example, comes out as one stem per speaker in the original turn order. In separation mode, the input audio-mix remains clean while the captions specify what to extract, so the model can decompose a recording from a free-text description. We evaluate generation on scenes synthesized from text, and separation on scenes rendered by other generators and on real recordings of speech, music, and sound effects. Given an audio-mix and captions that carry the spoken lines, SepGen outperforms language-conditioned separators, most clearly on speech, and it keeps the lead when the lines are removed from the captions. Code, checkpoints, and datasets are available at https://sepgen.github.io/
发表机构
- Tel Aviv University(特拉维夫大学)
- Google(谷歌公司)
机构由 AI 辅助整理,请以论文原文为准。