arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12250cs.SDcs.AI

嵌入中还保留了多少音频?音频编码器的反演审计

How Much Audio Is Left In An Embedding? An Inversion Audit Of Audio Encoders

Marios Glytsos, Brian McFee

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过配对源重建,利用 Stable Audio Open 解码器从多种音频编码器(VGGish、ConvNeXt、CLAP、EnCodec)的冻结表示中重建音乐,发现不同编码器的可重建性存在差异,压缩任务嵌入仍能保留源特异性与高级音乐内容。

中文摘要 AI 辅助

预训练音频编码器会被复用至下游任务,而这些下游任务在编码器训练时往往是未知的,因此其有用性部分取决于 pretext 目标保留了哪些信号属性。我们通过配对源重建来研究这种保留的信息。使用共享的 Stable Audio Open 潜在扩散解码器,我们从冻结表示中重建了5秒、44.1kHz的立体声音乐,这些冻结表示来自监督分类器(VGGish、ConvNeXt)、音频-文本对比模型(CLAP)和波形重建模型(EnCodec)。这些目标对源细节的保留施加了不同的压力,同时它们暴露的接口在时间和频谱分辨率上差异很大。在百万歌曲数据集(MSD)上进行评估,我们发现不同编码器系列的可重建性存在明显差异,而编码器内部的比较显示,当暴露更精细的时间或频谱结构时,恢复效果会提升。即使是压缩的面向任务的嵌入,也支持保留可测量的源特异性和高级音乐内容的重建。

英文摘要

Pretrained audio encoders are reused for downstream tasks that are often unknown when the encoder is trained, so their usefulness depends partly on which signal properties survive the pretext objective. We study this retained information through paired source reconstruction. Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec). These objectives impose different pressures to preserve source detail, while their exposed interfaces vary substantially in temporal and spectral resolution. Evaluating on the Million Song Dataset (MSD), we find clear differences in reconstructability across encoder families, while within encoder comparisons show improved recovery when finer temporal or spectral structure is exposed. Even compressed task oriented embeddings support reconstructions that preserve measurable source specificity and high level musical content.

发表机构

  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑