探索音频变换表征学习的设计空间
Exploring the Design Space of Representation Learning for Audio Transformations
浏览论文内容
中文总结 AI 辅助
该研究针对音频变换表征学习的设计选择问题,构建含三个目标的统一框架,揭示两类嵌入的互补作用,结合架构与训练改进后,表征在多项任务中优于现有基线。
中文摘要 AI 辅助
神经音频表征学习已支持一系列面向内容的应用,但生成的特征在涉及音频处理的任务中仍存在局限。此外,面向处理的表征应捕捉什么尚不明确:是处理本身(与源内容解耦),还是保留源内容的处理后音频。现有方法隐含地偏向其中一种,且在模型、数据和评估方面存在差异,导致难以明确哪些设计选择驱动了其性能。我们在包含三个目标(处理一致性、描述对齐及通过前向预测实现的等变性)的统一框架内解决上述两个问题。我们在受控设置下对比所有目标的组合,揭示其相对优势与相互作用。我们的框架同时生成变换嵌入和处理后音频嵌入,发现二者发挥互补作用:基于距离的任务更适合前者,而基于探测的任务更适合后者。结合网络架构和训练流程的改进,我们的表征在检索、基于探测的评估及风格迁移任务中均优于现有基线。
英文摘要
Neural audio representation learning has enabled a range of content-oriented applications, but the resulting features remain limited for tasks involving audio processing. Furthermore, it is not obvious what processing-aware representations should capture: the processing itself, abstracted away from source content, or the processed audio that retains it. Existing approaches implicitly commit to one or the other and also differ in their models, data, and evaluation, obscuring which design choices drive their behavior. We address both questions within a unified framework of three objectives: processing consistency, description alignment, and equivariance via forward prediction. We compare all combinations of the objectives under a controlled setup and reveal their relative strengths and interactions. Our framework produces both a transformation embedding and a processed-audio embedding, and we find that the two play complementary roles: distance-based tasks favor the former, while probe-based tasks favor the latter. Combined with improvements in network architecture and training pipeline, our representations outperform prior baselines across retrieval, probe-based evaluation, and style transfer.