AllMusicCaps:将专辑评论作为Music CLAP的补充监督
AllMusicCaps: Album Reviews as Complementary Supervision for Music CLAP
浏览论文内容
中文总结 AI 辅助
本研究提出AllMusicCaps,利用AllMusic专家专辑评论经LLM预处理构建标题语料库,结合SigReg正则化优化CLAP模型,提升了文本到音乐检索等任务性能并发布相关资源。
中文摘要 AI 辅助
近期的开放文本-音频对比模型(CLAP)通常使用从标签数据集或网页搜索结果中生成的LLM标题进行训练,这些标题往往准确但表达性较窄。作为补充来源,我们探索了人类撰写的专辑评论,特别是来自AllMusic的专家评论:这类评论规模庞大,且包含其他来源缺乏的叙事线索、评价性形容词和场景框架。由于原始评论噪声过大,无法直接用作标题,我们首先通过LLM预处理流程构建了包含245346个样本的标题语料库,该流程会识别描述性音乐引语并将其重写为可用于训练的标题。我们发现,专辑评论监督在人类撰写的标题基准(Song Describer)上带来了最大的检索增益,尤其是针对其他现有标题数据集未覆盖的复杂查询。此外,我们重新审视了训练方案,发现SigReg正则化(鼓励嵌入空间中的各向同性高斯分布)可改善分类任务中的MLP探测以及文本到音乐的检索。最终得到的模型在文本到音乐检索、零样本分类和大多数MLP探测任务上均优于开放CLAP风格的基线模型。我们发布了由评论衍生的标题数据集和模型权重,以支持未来的研究。
英文摘要
Recent open text-audio contrastive models (CLAPs) are typically trained with LLM-generated captions derived from tag datasets or web search results, which tend to be accurate but expressively narrow. As a complementary source, we explore human-written album reviews, specifically expert reviews from AllMusic: they exist at scale and carry narrative cues, evaluative adjectives, and scene framing that other sources lack. Since raw reviews are too noisy for direct use as captions, we first build a caption corpus with 24,5346 samples via an LLM preprocessing pipeline that identifies descriptive musical quotes and rewrites them into training-ready captions. We find that album review supervision yields the largest retrieval gains on a human-written caption benchmark (Song Describer), particularly for complex queries that other existing caption datasets leave uncovered. In addition, we revisit the training recipe and show that SigReg regularization, which encourages an isotropic Gaussian distribution in the embedding space, improves MLP probing across classification tasks, as well as text-to-music retrieval. The resulting model outperforms open CLAP-style baselines on text-to-music retrieval, zero-shot classification, and most MLP probing tasks. We release the review-derived caption dataset and model weights to support future research.