SAGE:语义音频生成编码器
SAGE: Semantic Audio Generative Encoder
浏览论文内容
中文总结 AI 辅助
SAGE是一种轻量级音频自编码器,通过蒸馏预训练音频-文本模型嵌入来塑造潜在空间,在重建质量、语义结构和推理速度间达到最佳平衡,并以更小规模超越更大模型。
中文摘要 AI 辅助
音频自编码器将波形压缩为紧凑的潜在表示,作为原始音频与下游模型之间的接口。现有系统在重建质量、潜在空间的语义结构和推理速度之间面临三向权衡,通常以牺牲其他方面为代价来优先考虑其中一两个方面。本文介绍了SAGE(语义音频生成编码器):一种紧凑的变分自编码器,仅使用公开可用的音乐进行训练,通过从预训练的音频-文本模型中提取嵌入来塑造其潜在空间。该105M参数模型以Stable Audio Open的推理成本运行,达到了SAME-L(一个8倍大且4倍慢的自编码器)的听感测试质量,同时在客观感知和分布重建指标上超越了这两者。此外,它在所有十九项潜在语义探测任务(包括领域内和领域外)上均达到了最先进水平。这些结果确立了SAGE作为一种轻量级音频自编码器的地位,在我们评估的模型中,它在三向权衡中达到了最佳平衡,结合了高重建保真度、最先进的语义结构和快速推理。
英文摘要
Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference speed, typically favoring one or two of these at the expense of the others. This paper introduces SAGE, Semantic Audio Generative Encoder: a compact variational autoencoder, trained solely on publicly available music, that shapes its latent by distilling embeddings from a pretrained audio-text model. This 105M-parameter model runs at the inference cost of Stable Audio Open and reaches the listening-test quality of SAME-L, an autoencoder 8x larger and 4x slower, while surpassing both on objective perceptual and distributional metrics of reconstruction. Furthermore, it sets the state of the art on all nineteen probing tasks of latent semantics, in domain and out of domain. These results establish SAGE as a lightweight audio autoencoder that strikes the best balance of the three-way trade-off among those we evaluate, combining high reconstruction fidelity, state-of-the-art semantic structure, and fast inference.