使用槽嵌入从音乐混合中检索单个音轨
Retrieving Individual Stems from Music Mixtures with Slot Embeddings
浏览论文内容
中文总结 AI 辅助
提出Stembed,利用槽嵌入将音乐混合编码为多个候选音轨表示,通过对比学习匹配独立音轨,在无族标签下超越CIR基线,实现更灵活的乐器检索。
中文摘要 AI 辅助
音乐制作人在乐器独立录音(称为音轨)的库中搜索与现有歌曲部分相似的声音。神经检索系统通过将音频映射到嵌入向量,并根据查询与库中音轨的相似度进行排序来解决这一问题。领先的方法,对比乐器检索(CIR),将混合音频编码为单个嵌入向量,但它在用户指定目标乐器族时效果最佳。我们提出了Stembed,它将混合音频编码为多个代表候选音轨的槽嵌入。在训练过程中,我们从同一首歌的音轨构建混合音频,并将其槽嵌入与独立音轨的嵌入进行匹配。来自混合音频的槽嵌入继承了其分配的独立嵌入的音轨身份,从而支持对比损失。在来自保留的MoisesDB艺术家的混合音频上,当两者都搜索整个音轨库时,Stembed优于CIR风格的基线。即使在无族标签的情况下预测音轨数量,Stembed也超过了基线的族过滤R@1。我们的网站演示了用户如何通过检查其检索音轨的标签来选择槽。
英文摘要
Music producers search libraries of isolated instrument recordings, called stems, for sounds resembling parts of an existing song. Neural retrieval systems address this by mapping audio to embeddings and ranking library stems by their similarity to the query. The leading method, Contrastive Instrument Retrieval (CIR), encodes the mixture as a single embedding, but it works best when a user specifies the target's instrument family. We introduce Stembed, which encodes a mixture as several slot embeddings representing candidate stems. During training, we construct mixtures from stems of the same song and match their slot embeddings to those of the isolated stems. The slot embeddings from mixtures inherit the stem identities of their assigned solo embedding, enabling a contrastive loss. On mixtures from held out MoisesDB artists, Stembed outperforms a CIR-style baseline when both search the full stem library. Even when predicting the stem count itself without family labels, Stembed exceeds the baseline's family-filtered R@1. Our website demonstrates how users can select a slot by inspecting the tags of its retrieved stems.
发表机构
- Princeton University(普林斯顿大学)
- The Ohio State University(俄亥俄州立大学)
- Symbal AI
机构由 AI 辅助整理,请以论文原文为准。