发表机构
Creative Computing Institute, University of the Arts London(伦敦艺术大学创意计算学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过逐层和跨层聚类分析RAVE与EnCodec解码器,发现合成及自然音频特征编码受输入、深度和分布影响,中间层联合编码更强,跨层聚类可提升BPM编码强度与普遍性,增强模型可解释性并指导神经合成控制。
AI 中文摘要
神经音频合成模型如实时音频变分自编码器(RAVE)实现了令人印象深刻的生成质量,但其内部表示如何编码音乐特征仍知之甚少。我们对RAVE解码器激活进行了系统的逐层和跨层聚类分析,涉及在三个不同音乐领域训练的模型,并使用四种刺激类型进行测试。随后,我们使用通用目的的EnCodec模型评估架构泛化能力。对于RAVE,我们发现合成刺激在模型和音频特征中编码良好(音高|ρ|=0.45,为零假设的5.1倍;BPM |ρ|=0.76,为零假设的8.6倍)。当使用自然音频时,这些结果有所降低但仍显著可见(特征平均|ρ|=0.25,为零假设的2.8倍)。自然音频在使用非线性探针时表现出更强的编码(特征平均R²=0.56,为零假设的18倍,非线性增益相对于线性探针R²为+0.152)。编码强度在解码器的各层中变化,且在所有音频特征中观察到中间层联合编码能力的增强(β2全部为负,p<0.05)。通用目的的EnCodec解码器在音频特征上也表现出类似的强合成响应,自然音频联合编码的类似非线性增益以及类似的深度分布。我们发现最佳跨层聚类相对于同一部分内的最佳整层,提高了BPM编码的强度(r=0.65,p=0.006)和普遍性(r=0.75,p=0.001),而对联合编码无影响。这些发现推进了神经音频模型的可解释性,并为神经合成的定向控制策略提供了信息。
英文摘要
Neural audio synthesis models like the Realtime Audio Variational autoEncoder (RAVE) achieve impressive genera tion quality, yet how their internal representations encode musical features remains poorly understood. We present a systematic layer-wise and cross-layer cluster analysis of RAVE decoder activations across three models trained on different musical domains, tested with four stimulus types. We then evaluate architectural generalization with a general purpose EnCodec model. For RAVE, we find that synthetic stimuli are encoded well across models and audio features (pitch |\r{ho}|=0.45, 5.1x the null, BPM |\r{ho}| = 0.76, 8.6x the null). These results are reduced but still substantively apparent when using natural audio (mean across features |\r{ho}|=0.25, 2.8x the null). Natural audio sees a stronger encoding when nonlinear probes are used (mean across features R2=0.56, 18x the null, +0.152 nonlinear gain over the linear probe R2). Encoding strength varies throughout the layers of the decoder and an increased ability to joint-encode in the middle layers is seen across all audio features (\b{eta}2 all negative, p < 0.05). The general purpose EnCodec decoder also sees similar strong synthetic responses across audio features, similar nonlinear gains for natural audio joint encoding and similar depth profiles. We find the best cross-layer cluster improves the strength (r = 0.65, p = 0.006) and prevalence (r = 0.75, p = 0.001) of BPM encoding when compared against the best whole layers within the same section, with no effect for joint encoding. These findings advance the interpretability of neural audio models and inform targeted control strategies for neural synthesis.
CommentsThis manuscript has been accepted for publishing in IEEE Transactions on Audio, Speech and Language Processing (TASLP)