Spot, Separate, and Enhance:全生成式音频混音方法
Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing
- Dolby Laboratories(杜比实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出首个多模态用户引导的生成模型SSE,通过视频和文本引导实现音频重混与增强,并构建DegradedMix数据集,实验证明其在可控性和重混质量上优于现有基线。
AI中文摘要:
我们提出了Spot, Separate, and Enhance (SSE),这是首个多模态、用户引导的音频重混与增强生成模型。SSE通过视频和文本描述的双重引导,重新平衡音频、移除不需要的音频源并减少混响,从而增强视频内容。为支持其训练与评估,我们提出了DegradedMix,这是一个基于音频重混基准MuddyMix构建的新数据集。我们还采用了生成建模中的评估指标,这些指标比标准的基于重建的指标更能捕捉重混的创造性本质。大量实验表明,SSE在可控性和重混质量方面均优于现有基线。项目页面:此https URL
英文摘要:
We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement. SSE enhances video content by rebalancing the audio, removing unwanted audio sources, and reducing reverberation, guided by both video and textual descriptions. To support its training and evaluation, we propose DegradedMix, a new dataset built on the audio remixing benchmark MuddyMix. We also adopt evaluation metrics from generative modeling, which better capture the creative nature of remixing than standard reconstruction-based metrics. SSE outperforms existing baselines in both controllability and remixing quality, as shown by extensive experiments. Project page: https://sse-ai.notion.site