发表机构
Imperial College London; University of Surrey(帝国理工学院; 萨里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出自监督框架NAPE,通过因果Transformer预测下一个音频补丁嵌入,在六个音频语音基准上实现先进性能,可扩展编码器规模,获强线性探测结果,无显式监督时生成结构化注意力模式。
AI 中文摘要
自监督学习(SSL)已推动音频表示学习取得重大进展,不过现有方法日益依赖复杂的预训练方案来达到有竞争力的性能。语言建模领域以及近期视觉表示学习领域最具影响力的进展背后,支撑着一种截然不同的预训练理念:与其将编码器训练为静态特征提取器,不如训练模型从先前的上下文预测下一个元素——离散令牌或连续嵌入。自回归预测因此提供了一个跨模态迁移的统一预训练接口,促使模型学习底层数据分布。我们探究,鉴于音频的时间结构使补丁嵌入的自回归预测成为自然适配,这种简单的因果范式能否产生强大的音频学习器。我们引入NAPE(下一个音频补丁嵌入预测),这是一个自监督框架,其中因果Transformer通过因果掩码和停止梯度作为唯一训练信号,从先前的对数梅尔频谱图补丁嵌入中预测每个下一个补丁嵌入。该设计刻意极简,避免了重建解码器、声学令牌生成器、师生设置以及辅助正则化损失。在六个音频和语音基准测试中,NAPE在多项任务上实现了最先进的微调性能,在编码器规模上可一致扩展,并产生强大的线性探测结果。NAPE还在无显式监督的情况下生成结构化注意力模式。
英文摘要
Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing methods have increasingly relied on elaborate pre-training recipes to reach competitive performance. A markedly different pre-training philosophy underpins the most influential progress in language modeling and, more recently, in visual representation learning: rather than train encoders as static feature extractors, models are trained to predict the next element, a discrete token or a continuous embedding, from the preceding context. Autoregressive prediction thereby provides a unified pre-training interface that transfers across modalities, compelling the model to learn the underlying data distribution. We ask whether such a simple causal paradigm can yield strong audio learners, given that audio's temporal structure makes autoregressive prediction of patch embeddings a natural fit. We introduce NAPE (Next-Audio-Patch-Embedding prediction), a self-supervised framework in which a causal Transformer predicts each next patch embedding of a log-mel spectrogram from the previous ones, using causal masking and stop-gradient as its sole training signal. The design is intentionally minimalist, avoiding reconstruction decoders, acoustic tokenizers, student-teacher setups, and auxiliary regularization losses. Across six audio and speech benchmarks, NAPE achieves state-of-the-art fine-tuning performance on several tasks, scales consistently across encoder sizes, and yields strong linear-probing results. NAPE also produces structured attention patterns without explicit supervision.
CommentsProject website: https://umbertocappellazzo.github.io/nape