arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mizar:一个用于音频理解的1.59亿参数音频语言模型

Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding

Kaiyang Li, Shaobo Han, Yue Tian, Shihao Ji

arXiv 2609.28344首次发表:更新:

发表机构

NEC Labs(NEC实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Mizar,一个1.593亿参数的音频语言模型,通过架构、数据和三阶段训练,在三个基准上超越同类模型,并支持单CPU低延迟推理。

AI 中文摘要

音频语言模型(ALMs)将声学感知与语言模型中编码的知识相结合,从而能够对听觉事件进行情境化理解。使得这些能力在内存和计算资源有限的设备上变得实用,促使我们关注参数少于2亿的小型ALMs。我们提出了一种结合架构、数据和三阶段训练的配方,以构建Mizar,一个拥有1.593亿参数的ALM。其架构通过一个频率合并映射器将紧凑的CED-Small音频编码器连接到SmolLM2-135M。在来自ReasonAQA、AudioMCQ和AVQA的监督下,模型经历三个训练阶段:音频-语言对齐(阶段1)、音频相关微调(阶段2)和后训练(阶段3),旨在加强薄弱技能同时保留已学能力。在五个随机种子上,Mizar在MMAU上达到52.92%的平均准确率,在MMAR上达到42.42%,在ADQA-clean上达到36.02%,在三个基准上均超过了之前性能最佳的低于2亿参数的ALM。它还支持在单个CPU上进行本地推理:对于MMAU基准中的问题,从打开音频文件到生成完整答案的平均延迟为1.09秒。代码和检查点可在https://github.com/KaiyangLi1992/Mizar_159M获取。

英文摘要

Audio-language models (ALMs) integrate acoustic perception with the knowledge encoded in language models, enabling contextual understanding of auditory events. Making these capabilities practical on devices with limited memory and computation motivates our focus on small ALMs with fewer than 200M parameters. We introduce a recipe that brings together architecture, data, and three-stage training to build Mizar, a 159.3M-parameter ALM. Its architecture connects a compact CED-Small audio encoder to SmolLM2-135M through a frequency-merging mapper. With supervision drawn from ReasonAQA, AudioMCQ, and AVQA, the model undergoes three training stages: audio-language alignment (Stage 1), audio-dependent fine-tuning (Stage 2), and post-training (Stage 3) aimed at strengthening weak skills while retaining learned capabilities. Across five random seeds, Mizar achieves mean accuracies of 52.92% on MMAU, 42.42% on MMAR, and 36.02% on ADQA-clean, surpassing the previous best-performing ALM below 200M parameters on all three benchmarks. It also supports local inference on a single CPU: on questions from the MMAU benchmark, the mean latency from opening the audio file to generating a complete answer is 1.09 seconds. Code and checkpoints are available at https://github.com/KaiyangLi1992/Mizar_159M.

Comments5 pages, submitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑