arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SALMONN-2:利用自监督表示提升通用听力能力

SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations

Xiaoyu Yang, Xuenan Xu, Wenyi Yu, Siyin Wang, Changli Tang, Terumi Chiba, Siyuan Hou, Ziyang Zhang, Wen Wu, Baoxiang Li, Guangzhi Sun, Chao Zhang, Philip Woodland

arXiv 2607.17079首次发表:更新:

AI 中文总结

研究探索自监督学习音频表示能否为音频大语言模型提供基础,提出基于统一SSL编码器的SALMONN-2,用多层特征融合适配器,还探索多模态上下文学习。实验表明通用SSL编码器性能良好,SALMONN-2达先进水平,且特定训练可获多模态上下文学习能力。

AI 中文摘要

近期音频大语言模型(ALLMs)通常基于用大量监督数据训练的音频编码器构建。鉴于自监督学习(SSL)音频编码器模型能学习通用且可转移的表示,研究其能否为ALLMs提供有效基础。提出基于统一SSL编码器的ALLM——SALMONN-2,为更好利用SSL编码器学习的分层表示,提出多层特征融合(MLF)适配器。还探索ALLMs中的多模态上下文学习(MICL)及获取该能力的方法。实验表明通用SSL编码器性能与专用监督音频编码器相当或更优,SALMONN-2在ALLM理解基准测试中达先进性能,且MICL可通过针对性上下文偏差训练有效获得。

英文摘要

Recent audio large language models (ALLMs) are typically built upon audio encoders trained with large amounts of supervised data. Since self-supervised learning (SSL) audio encoder models are known to learn general-purpose and transferable representations, we investigate whether general-purpose SSL audio representations can serve as an effective foundation for ALLMs. We present SALMONN-2, an ALLM built upon a unified SSL encoder. To better exploit the hierarchical representations learned by SSL encoders, we propose a multi-layer feature fusion (MLF) adapter that aggregates information from all encoder layers before projecting them into the language model. Beyond conventional audio understanding tasks, we further explore multimodal in-context learning (MICL) in ALLMs and study how this capability can be acquired through contextual biasing training. Experimental results show that a general-purpose SSL encoder achieves performance comparable to, or better than, specialised supervised audio encoders while providing a more balanced capability across speech, audio, music and paralinguistic tasks. SALMONN-2 further achieves state-of-the-art performance among comparable-scale open-weight models on ALLM understanding benchmarks, obtaining the best results on MMAU-Pro, MMAR and MMSU. We also show that MICL does not emerge naturally in ALLMs, but can be effectively acquired through targeted contextual biasing training.

CommentsDisclaimer: This work has been submitted to the IEEE for possible publication

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑