发表机构
Harbin Institute of Technology(哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SRF-SVB,利用修正流实现高保真、高效的歌声美化,通过上下文引导的掩码梅尔频谱图修补保留歌手风格,在英中测试集上优于基线。
AI 中文摘要
歌声美化(SVB)旨在纠正业余歌唱的音高和节奏,同时提升嗓音质量,并保留歌词和歌手的音色。然而,现有方法在生成质量和效率上存在局限,且往往忽视对歌手风格的保留。我们提出SRF-SVB,一种基于修正流的风格一致歌声美化模型,实现了涵盖音高和节奏校正的高保真、高效美化。此外,我们设计了一种上下文引导的掩码梅尔频谱图修补机制,有效保留了业余歌手的风格,包括独特音色和表现模式。在英文和中文测试集上的实验表明,SRF-SVB在大多数客观和主观指标上优于基线模型。
英文摘要
Singing voice beautifying (SVB) aims to correct pitch and rhythm of amateur singing while enhancing vocal quality, preserving lyrics and the singer's timbre. Existing methods, however, suffer from limited generation quality and efficiency, and tend to neglect the preservation of the singer's style. We propose SRF-SVB, a style-consistent model for SVB via rectified flow, which achieves high-fidelity and efficient beautification covering pitch and rhythm correction. Furthermore, we design a context-guided masked mel-spectrogram inpainting mechanism that effectively preserves the amateur singer's style, including unique timbre and expressive patterns. Experiments on both English and Chinese test sets show that SRF-SVB outperforms baseline models in most objective and subjective metrics.