发表机构
Awarri Technologies Limited; GSM Association(Awarri科技有限公司; GSM协会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该工作提出离线语音优先AI架构及硬件参考栈,通过量化与基准测试验证Q4_K_M为最佳权衡,在边缘设备上实现高效多语言部署。
AI 中文摘要
离线AI模块工作流使得完全离线运行的语音优先AI系统能够实现实用、低功耗且社区可访问的部署。该工作流专为非洲语言社区设计,在这些社区中,语音是主要的交互方式,而互联网连接不可靠或不存在。该工作流提供了三个相互加强的组件:一个模块化的语音优先离线架构、一个低成本的硬件参考物料清单,以及一个针对2-5B参数类指令调优语言模型的可复现量化与基准测试流水线。本文首次对两个硬件层级上的整个栈进行了端到端基准评估:NVIDIA Jetson Orin NX(TierB)和Raspberry Pi5(TierA)。三个指令调优模型在四种量化格式下进行评估,评估内容包括部署指标(解码吞吐量、聊天延迟、内存、功耗)和多语言质量(在MasakhaNEWS上对英语、豪萨语、伊博语、尼日利亚皮钦语和约鲁巴语的主题分类准确率;每种语言的困惑度漂移)。语音识别使用Ethio-ASR在阿姆哈拉语和奥罗莫语上对两个层级进行评估。主要发现是Q4_K_M量化代表了在两个层级上部署的最佳大小-质量权衡:在TierB上,gemma-4-E2B-it在Q4_K_M下实现了28.8t/s的解码吞吐量和89.2%的主题分类准确率,而所有三个模型在TierA上均在16GB内存预算内运行。
英文摘要
The Offline AI Modules workstream enables practical, low-power, and community-accessible deployment of voice-first AI systems that operate fully offline. Designed for African language communities where speech is the dominant mode of interaction and internet connectivity is unreliable or absent, the workstream delivers three reinforcing components: a modular voice-first offline architecture, a low-cost hardware reference bill of materials, and a reproducible quantization and a reproducible quantization and benchmarking pipeline for instruction-tuned language models in the 2-5B parameter class. This paper presents the first end-to-end benchmark evaluation of the stack across two hardware tiers: an NVIDIA Jetson Orin NX (TierB) and a Raspberry Pi5 (TierA). Three instruction-tuned models are evaluated across four quantization formats, assessed for deployment metrics (decode throughput, chat latency, memory, power) and multilingual quality (topic classification accuracy on MasakhaNEWS across English, Hausa, Igbo, Nigerian Pidgin, and Yoruba; per-language perplexity drift). Speech recognition is evaluated using Ethio-ASR on Amharic and Oromo across both tiers. The principal finding is that Q4_K_M quantization represents the best size-to-quality trade-off for deployment on both tiers: gemma-4-E2B-it achieves 28.8t/s decode throughput and 89.2% topic classification accuracy at Q4_K_M on TierB, while all three models run within the 16GB memory budget on TierA.