不要 CLAP:音乐-文本模型是词袋模型吗?
Don't CLAP: Are Music-Text Models Bag-of-Words?
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过属性交换扰动测试,发现CLAP等音乐-文本模型无法区分改变含义的标题扰动,其表示接近词袋,缺乏细粒度音乐语义绑定能力。
AI中文摘要:
文本到音乐系统根据音频质量和音乐遵循提示的忠实程度进行评估,而CLAP分数,即音乐-文本模型的音频与文本嵌入之间的余弦相似度,是衡量忠实度的标准客观指标。我们探究该分数在多大程度上准确反映文本:当一个属性与乐器相关联(如失真吉他)时,文本嵌入是否捕捉到这种绑定?为了查明,我们引入一种属性交换扰动:通过交换两个乐器之间的恰好一个属性(音色、主奏与伴奏,或首次出现的顺序)来编辑真实录音的标题。然后,我们测试四个对比音乐-文本模型和一个大型音频-语言模型,看音频对原始标题的评分是否高于对扰动标题的评分。没有一个对比模型能可靠地区分这两个标题。音频-语言模型表现更好,但进一步实验表明,其优势在很大程度上依赖于与音频无关的语言先验。因此,我们的结果提供了令人信服的证据,表明CLAP分数及相关指标并未捕捉到细粒度的音乐含义或属性绑定;它们的表示更接近词袋,使其对标题中改变含义的扰动不敏感。
英文摘要:
Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask how accurately that score reflects the text: when an attribute is linked to an instrument (e.g., distorted guitar), does the text embedding capture that binding? To find out, we introduce an attribute swap perturbation: the caption of a real recording is edited by exchanging exactly one property, timbre, lead versus accompaniment, or order of first appearance, between two instruments. We then test four contrastive music-text models and one large audio-language model on whether the audio scores higher against the original caption than against the perturbed one. No contrastive model distinguishes the two captions reliably. The audio-language model does better, but further experiments show that its advantage rests largely on audio-agnostic language priors. Our results thus provide compelling evidence that the CLAP score and related metrics do not capture fine-grained musical meaning or attribute bindings; their representation is closer to a bag-of-words that leaves them insensitive to meaning-changing perturbations of the caption.