发表机构
Georgia Institute of Technology; University of Rochester(佐治亚理工学院; 罗切斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MuseTimbre是首个通过调节预训练音乐生成器,从音频参考向复音声源迁移音色的零样本系统,利用多音高估计和微调CLAP编码器,在四个数据集上实现与基线相当的音高对齐和更紧密的音色匹配。
AI 中文摘要
乐器音色迁移利用另一种乐器的音色重新演绎一段演奏。从音频参考中提取目标音色比从文本提示中推断能捕捉更多细微差别。从这类片段中读取音色的系统会为任务训练专用模型,这能干净地捕捉音色,但仍是狭窄的、单一用途的系统。更通用的方法是为预训练的音乐生成器添加控制,然而参考片段将音色与流派和旋律纠缠在一起,因此这些系统退而求其次,用文本来命名音色。我们提出MuseTimbre,据我们所知,这是首个通过调节预训练音乐生成器,将音色从音频参考迁移到复音声源的系统。该系统采用多音高估计器从声源提取音高信息,并微调CLAP编码器以从参考音频中提取音色信息。实验表明,在四个真实复音录音数据集上,MuseTimbre在音高对齐方面与基线相当,同时更紧密地匹配参考音色。结果还表明,微调后的基于CLAP的音色提取器对音高变化具有鲁棒性,使其在音色相似度度量中很有用。
英文摘要
Instrument timbre transfer re-voices a performance using the timbre of another instrument. Extracting the target timbre from an audio reference capture more nuances than inferring it from a text prompt. Systems that read timbre from such a clip train a dedicated model for the task, which captures the timbre cleanly but stays a narrow, single-purpose system. More versatile approaches add control to a pretrained music generator, yet a reference clip entangles timbre with genre and melody, so these systems fall back on text to name the timbre. We present MuseTimbre, the first system, to our best knowledge, that transfers timbre from an audio reference to a polyphonic source through conditioning a pretrained music generator. This system employs a multi-pitch estimator to extract pitch information from the source and finetune a CLAP encoder to extract timbre information from the reference audio. Experiments show that across four datasets of real polyphonic recordings, MuseTimbre achieves pitch alignment on par with the baselines while matching the reference timbre far more closely. Results also show that the finetuned CLAP-based timbre extractor is robust to pitch variations, making it useful in timbre similarity measures.
Comments4 pages plus references, 4 figures, 2 tables