arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CrowdioSet与PaRIRset:面向现场音乐源分离的两个数据集

CrowdioSet and PaRIRset: Two Datasets Towards Live Music Source Separation

Enric Gusó, Xavier Serra

arXiv 2607.27828首次发表:更新:

AI 中文总结

该研究针对现有音乐源分离模型难以泛化到现场音乐的问题,构建CrowdioSet与PaRIRset两个数据集,可提升模型现场分离性能,相关资源已公开。

AI 中文摘要

大多数音乐源分离(Music Source Separation, MSS)模型仅在录音室录音上训练,无法很好地泛化到现场音乐录音,因为它们忽略了场馆声学、扬声器系统响应以及观众噪声。我们通过提供两个新数据集并在其上训练模型来弥合这一差距。首先,我们推出CrowdioSet:一个噪声数据集,包含来自Freesound的4800条真实环境音轨,以及从MUSDB18和MOISESDB数据集人声经零样本歌声转换生成的合成跟唱声。CrowdioSet可对现场录音进行有效音频去噪,在客观和主观评估中均实现更优的分离效果。其次,我们引入PaRIRset:一个使用麦克风阵列在40个专业演唱会场馆采集的立体声脉冲响应数据集。我们的结果表明,与仅使用语音增强任务的真实脉冲响应相比,添加PaRIRset的脉冲响应可提升MSS模型的性能。我们向公众免费提供示例、代码、模型权重、PaRIRset和CrowdioSet。

英文摘要

Most Music Source Separation (MSS) models do not generalize well to live music recordings because they are trained on studio recordings alone, disregarding the venue acoustics, the speaker system's response and audience noise. We propose to bridge this gap by providing and training a model on two novel datasets. First, we present CrowdioSet: a noise dataset comprising 4800 real ambience tracks from Freesound and synthetic sing-alongs for the vocals in MUSDB18 and MOISESDB datasets, generated from zero-shot singing voice conversions. CrowdioSet enables effective audio denoising for live recordings, resulting in superior separation both in objective and subjective evaluations. Second, we introduce PaRIRset, a stereo impulse response dataset captured across 40 professional concert venues using a microphone array. Our results show that adding PaRIRset RIRs increases the performance of a MSS model compared to using real RIRs from Speech Enhancement tasks alone. We make the examples, code, model weights, PaRIRset, and CrowdioSet freely available to the public.

CommentsAccepted to ISMIR26. See : https://enricguso.github.io/crowdioset_parirset

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑