音频表示是否可加性组合?
Do Audio Representations Compose Additively?
浏览论文内容
中文总结 AI 辅助
本研究通过线性对齐和留一组合外推诊断,探究预训练音频表示是否内化加性组合结构,发现CLAP优于基线,但所有模型均存在重建残差,表明加性组合性有限。
中文摘要 AI 辅助
组合性,即将复杂声学场景表示为更简单声源组合的能力,是听觉感知和经典加性信号模型的核心。然而,现代预训练音频表示在无组合监督的情况下是否内化了加性结构,仍不清楚。现有的音频组合推理评估框架主要关注跨模态音频-文本对齐,而未解决音频表示本身是否展现出独立于文本基础的加性组合结构,类似于词表示中的向量算术。为探究此问题,我们对冻结音频表示采用两步诊断法。首先,使用典型相关分析量化表示与声源标签之间的线性对齐。其次,通过留一组合外推测试加性组合泛化能力,即按精确声源标签集对片段分组,平均其表示,并仅基于训练组合拟合的每声源贡献来预测留出组的均值。在更大的组合留出情况下,CLAP在FSD50K上优于置换基线和标签重叠基线,而言语模型未优于标签重叠基线。我们在FSD50K和CHiME-Home数据集上检验了Wav2Vec2、HuBERT和CLAP生成的表示。所有三个模型均表现出比置换基线更高的线性相关性和更准确的留一组合外推重建。然而,仅CLAP显示出较大的余弦相似度增益,这可能与其在多种音频和文本上的训练有关。最后,我们注意到所有三个模型均存在重建残差,揭示了加性组合性的局限性,如非线性或非组合性的音频结构。
英文摘要
Compositionality, the ability to represent complex acoustic scenes as combinations of simpler sound sources, is central to auditory perception and classical additive signal models. Still, it remains unclear whether modern pre-trained audio representations internalize additive structure without compositional supervision. Existing evaluation frameworks of audio compositional reasoning largely focus on cross-modal audio-text alignment, leaving open whether audio representations themselves exhibit additive compositional structure independent of text grounding, analogous to vector arithmetic in word representations. To investigate this, we adopt a two-step diagnostic for frozen audio representations. First, we quantify linear alignment between representations and sound source labels using canonical correlation analysis. Second, we test additive compositional generalization via leave-one-combination-out reconstruction, grouping clips by exact source-label set, averaging their representations, and predicting held-out means from per-source contributions fitted only on training combinations. With larger combination holdouts, CLAP outperforms the permuted and label-overlap baselines on FSD50K, while the speech models do not outperform the label-overlap baseline. We examine representations generated by Wav2Vec2, HuBERT, and CLAP on FSD50K and CHiME-Home datasets. All three models show consistently higher linear correlation and more accurate leave-one-combination-out reconstructions than the permuted baselines. However, only CLAP shows large cosine similarity gains, which could be associated with its training on many kinds of audio and text. Finally, we note that all three models exhibit reconstruction residuals, revealing limits of additive compositionality such as nonlinear or non-compositional audio structure.
发表机构
- University of Oxford(牛津大学)
- University of Essex(埃塞克斯大学)
机构由 AI 辅助整理,请以论文原文为准。