AI 中文总结
针对歌声分离训练数据有限问题,将其作为从语音增强到歌声分离的域适应,研究全量微调和LoRA参数高效微调策略,实验表明两种策略均有效提升歌声分离性能,LoRA微调还能保留语音增强能力,证明适配预训练模型是数据稀缺时训练歌声分离模型的有效策略。
AI 中文摘要
先进的语音增强模型受益于大规模标记数据集,而歌声分离模型的可用训练数据有限。为解决这一限制,我们将歌声分离表述为从语音增强到歌声分离的域适应。我们研究了两种微调策略:全量微调以及在判别模型和生成模型上使用低秩适应(LoRA)的参数高效微调。采用任一适应策略的模型在信号失真比上比从头训练的相同架构高出0.29 - 1.8 dB。全量微调产生最高的歌声分离性能,但灾难性遗忘会降低语音增强性能。LoRA微调在仅比基础语音增强模型增加6 - 12%参数的情况下,实现了有竞争力的歌声分离性能,同时保留了原始语音增强能力。此外,生成模型对未见测试集的泛化能力有所提升。结果表明,在数据稀缺场景下,适配预训练的语音增强模型是训练歌声分离模型的有效策略。
英文摘要
State-of-the-art speech enhancement models benefit from large-scale labeled datasets, whereas singing voice separation models suffer from limited available training data. To address this limitation, we formulate singing voice separation as domain adaptation from speech enhancement to singing voice separation. We investigate two fine-tuning strategies: full fine-tuning and parameter-efficient fine-tuning using Low-Rank Adaptation (LoRA) on a discriminative and a generative model. Models with either adaptation strategy outperform the same architectures trained from scratch by 0.29-1.8 dB in Signal-to-Distortion-Ratio. Full fine-tuning yields the highest singing voice separation performance, but catastrophic forgetting degrades speech enhancement performance. LoRA fine-tuning achieves competitive singing voice separation performance while preserving the original speech enhancement capability with only 6-12% additional parameters compared to the base speech enhancement model. Furthermore, the generative model shows improved generalization to an unseen test set. The results demonstrate that adapting pretrained speech enhancement models is an effective strategy for training singing voice separation models in data-scarce scenarios.
CommentsAccepted for presentation at the International Workshop on Acoustic Signal Enhancement (IWAENC) 2026