arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniAE-MoE:一种基于混合专家模型的统一音频编码器

UniAE-MoE: A Unified Audio Encoder via Mixture of Experts

Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, Zhiyong Wu

arXiv 2609.39199首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于混合专家架构的统一音频编码器UniAE-MoE,融合多个主流编码器并采用两阶段微调与数据缩放技术,在基准上达到最先进性能。

AI 中文摘要

大型音频语言模型(LALMs)依赖有效的音频编码器来实现多任务性能。我们引入了UniAE-MoE,一种统一的音频编码器,旨在通过混合专家(MoE)架构建模跨域音频表示,并实现出色的下游理解性能。具体而言,我们探索了主流音频编码器,并整合了来自Qwen2-Audio和Audio-Flamingo 3的编码器,这些编码器展现出优越的下游能力。为了促进有效的模型融合,我们使用带有共享专家的SwiGLU来改进编码器,以解耦编码器网络,并进一步引入两阶段指令微调策略,以更好地使模型适应多样化的下游任务。此外,我们提出了任务特定数据缩放(TSDS)技术,以增强工具的理解能力。在XARES-LLM基准上,UniAE-MoE取得了0.802的分数,达到了最先进的性能。它还在官方Interspeech 2026音频编码器能力挑战赛中提供了顶级性能,进一步展示了在多样化音频任务中的强大泛化能力。这些结果共同验证了该工具在语音、音乐和通用音频领域的统一音频理解方面的有效性。

英文摘要

Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Experts (MoE) architecture. Specifically, we explore mainstream audio encoders and integrate those from Qwen2-Audio and Audio-Flamingo 3, which demonstrate superior downstream capabilities. To facilitate effective model fusion, we improve our encoder using SwiGLU with shared experts to decouple encoder networks, and we further introduce a two-stage instruction-tuning strategy to better adapt the model to diverse downstream tasks. Moreover, we propose the task-specific data scaling (TSDS) technique to enhance UniAE-MoE's understanding capabilities. On the XARES-LLM benchmark, UniAE-MoE attains a score of 0.802, achieving state-of-the-art performance. It also delivers top-tier performance in the official Interspeech 2026 Audio Encoder Capability Challenge, further demonstrating robust generalization across diverse audio tasks. Together, these results validate the effectiveness of UniAE-MoE for unified audio understanding across speech, music, and general audio domains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑