发表机构
Eximius Labs; Wabash College(埃克西米厄斯实验室; 沃巴什学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现对比音频嵌入模型编码的属性由语料库结构而非规模决定,通过调整语料标注多样性可提升情感识别性能,同时平衡关键词检测能力。
AI 中文摘要
当对比表征缺少某一属性时,扩大语料规模是默认的解决方案。本文报告了一个该方案失效的案例,并确定了真正有效的方法:在冻结的多模态嵌入模型基础上增加词汇-语音轮次,可将零样本关键词检测提升76个百分点,同时将语音情感识别降低14个百分点。该损失并非容量限制:对来自韵律控制语料库的7442个片段进行微调,可在付出5个百分点关键词检测代价的情况下,将情感识别性能恢复至语音前水平。这也与数据量无关:29428个标注明确提及情感的挖掘片段,在匹配曝光度下仅使情感识别性能提升-0.0007。差异源于结构:对比目标仅在批次内负样本无法被分离时才会编码某一属性;控制语料库固定句子内容,因此韵律是唯一可分离信号,而挖掘的标注虽提及情感,但仍可通过场景内容分离。对同一音频的干预验证了因果关系:提升标注相似度无法恢复情感,但若压缩标注多样性使情感成为唯一分离轴,可在三个随机种子下将情感识别性能提升8.9个百分点,在非表演语料库上也有小幅同符号增益,同时关键词准确率会反向变化。语料库结构而非规模或标注词汇,决定了对比音频嵌入模型编码的属性。
英文摘要
Scaling the corpus is the default remedy when a contrastive representation lacks an attribute. We report a case where it does nothing, and identify what does: adding a lexical-speech round to a frozen-base multimodal embedding model raises zero-shot keyword spotting by 76 points while reducing speech-emotion recognition by 14. The loss is not a capacity limit: fine-tuning on 7,442 clips from a prosody-controlled corpus recovers emotion past its pre-speech level at a five-point keyword cost. Nor is it data volume: 29,428 mined clips whose captions explicitly name emotions, at matched exposure, move emotion by -0.0007. The difference is structural: a contrastive objective encodes an attribute only when the in-batch negatives cannot be separated without it; the controlled corpus holds sentence content fixed, so prosody is the only separating signal, whereas mined captions name emotion yet remain separable by scene content. Intervention on the same audio confirms causality: raising caption similarity does not recover emotion, but collapsing caption diversity so that emotion becomes the only separating axis recovers it by 8.9 points across three seeds, with a smaller, same-signed gain on a non-acted corpus, while keyword accuracy trades back. Corpus structure, not size or caption vocabulary, controls what a contrastive audio embedding encodes.
Comments10 pages, 4 figures