Samsone:面向设备端推理的开放小型音频语言模型系列
Samsone: A Family of Open Small Audio Language Models for On-Device Inference
浏览论文内容
中文总结 AI 辅助
本文提出Samsone系列小型音频语言模型,面向设备端推理,在保持紧凑体积的同时达到与大规模模型竞争的性能,并开源代码、权重及Android应用。
中文摘要 AI 辅助
大型音频语言模型的成功推动了参数超过数十亿的庞大多模态网络的发展。然而,对隐私保护和低延迟处理的需求已将焦点转向能够在设备端执行的小型音频语言模型(SALMs)。在本文中,我们介绍了Samsone,一个专为边缘计算设计的小型音频语言模型系列。我们的核心模型Samsone-134M在多个基准测试中为其规模级别树立了新的最先进水平。我们进一步通过引入Samsone-99M和Samsone-356M来探索SALMs的缩放定律。尽管体积紧凑,Samsone系列的性能可与规模大几个数量级的模型相媲美。为了促进开放研究和可复现性,我们在公开可用的数据上训练Samsone。我们发布了训练代码、模型权重、移动端优化检查点,并提供了一个开源的Android应用程序,以演示Samsone的实时设备端推理。
英文摘要
The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacy-preserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-device execution. In this paper, we introduce Samsone, a family of SALMs designed for edge computing. Our core model, Samsone-134M, establishes a new state-of-the-art for its size class across multiple benchmarks. We further explore the scaling laws of SALMs by introducing Samsone-99M and Samsone-356M. Despite their compact footprint, the Samsone family delivers performance competitive with models orders of magnitude larger. To foster open research and reproducibility, we train Samsone on publicly available data. We release the training code, model weights, mobile-optimized checkpoints and provide an open-source Android application to demonstrate real-time on-device inference of Samsone.
发表机构
- AGH University of Kraków(克拉科夫AGH大学)
机构由 AI 辅助整理,请以论文原文为准。