arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩展音乐标注模式:零样本预测还是少样本适应?

Extending Music Annotation Schemas: Zero-Shot Prediction or Few-Shot Adaptation?

Christos Plachouras, Emmanouil Benetos, Johan Pauwels

arXiv 2610.06920首次发表:更新:

发表机构

Queen Mary University of London(伦敦玛丽女王大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对音乐标注模式扩展问题,提出基于MGPHot数据集的基准,对比零样本预测与监督适应,发现监督适应在较小预算下更优,冻结表示复用为适度预算下最有效方案。

AI 中文摘要

自动音乐标注通常在固定标注模式的假设下进行。然而在实践中,商业音乐目录往往需要随着需求的变化而容纳新的音乐属性。鉴于专家音乐标注成本高昂,目前尚不清楚哪种方法论在容纳新属性和回填现有曲目方面最为有效;音频-语言模型承诺实现零样本预测,但在何种标注预算下,监督适应会变得更具吸引力?我们基于MGPHot流行音乐标注数据集提出了一个基准,用于模拟不同标注预算下的音乐模式扩展。我们研究了使用音频-语言模型进行零样本预测、从预训练表示中学习新属性,以及适应在现有标注上训练的模型。我们的结果表明,即使标注预算很小,监督适应也比零样本预测更有效,而冻结表示复用对于适度的预算仍然是最有效的方法,无需深度适应所需的调参。

英文摘要

Automatic music annotation is typically tackled under the assumption of a fixed annotation schema. In practice, commercial music catalogs often need to accommodate new musical attributes as needs evolve. Given that expert music annotation is expensive, it is not evident which methodological approach is most effective at accommodating new attributes and backfilling existing tracks; audio-language models promise zero-shot prediction, but at what annotation budget does supervised adaptation become more compelling? We propose a benchmark based on the MGPHot popular music annotation dataset for simulating music schema extension across different annotation budgets. We investigate zero-shot prediction with audio-language models, learning new attributes from pretrained representations, and adapting models trained on existing annotations. Our results suggest that supervised adaptation is more effective than zero-shot prediction even with small annotation budgets, while frozen representation reuse remains the most effective approach for modest budgets without the tuning required by deeper adaptation.

Comments5 pages, 3 figures. Submitted to IEEE ICASSP 2027; under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑