arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31511cs.CL

Muslim:一个面向扎根伊斯兰知识的已部署阿拉伯语语音AI平台

Muslim: A Deployed Arabic Voice AI Platform for Grounded Islamic Knowledge

Yahya Mohamed Elnawasany

AI总结:

本文介绍Muslim,一个生产级阿拉伯语语音AI平台,提供有来源的伊斯兰知识,通过实时语音流水线、多源检索层、微调模型及计量与可观测性设计,实现高准确率(98.4%)和低延迟(0.9-1.7秒),并讨论了工程权衡。

AI中文摘要:

我们介绍Muslim,一个生产级的阿拉伯语语音AI平台,为真实用户提供有依据、有来源的伊斯兰知识。除了实时语音流水线(NeMo阿拉伯语语音识别、兼容OpenAI的LLM端点、自托管文本转语音)以及一个跨六个Model Context Protocol服务器路由的确定性多源检索层之外,我们报告了研究原型通常缺乏的三项内容。首先,发布了一系列微调的阿拉伯语伊斯兰模型工件:一个高效的工具路由LLM(Muslim-6B-PRO,5.94B参数)和一个现代标准阿拉伯语TTS模型(Fasih-TTS-V1),该模型在社区投票的阿拉伯语TTS竞技场(针对MSA)中排名第17名中的第5名,并在11个开放权重系统中排名第2。其次,一个账户和计量层——每个账户的免费回合配额、基于容量的拒绝以及延迟到真正需要时才进行的电子邮件验证——将开放演示转变为一个可操作、抗滥用的产品。第三,一个三层可观测性栈(存活检查、错误报告、产品分析),专门围绕系统的典型故障模式构建:GPU绑定的代理主机静默,而Web层继续正常服务。我们报告了真实测量的延迟和准确率数据(在124个案例上达到98.4%的背诵验证准确率;端到端语音延迟为0.9-1.7秒),并讨论了在生产中运行伊斯兰知识语音产品的具体工程权衡和局限性。

英文摘要:

We present Muslim, a production Arabic voice AI platform serving grounded, sourced Islamic knowledge to real users. Beyond a real-time voice pipeline (NeMo Arabic ASR, an OpenAI-compatible LLM endpoint, self-hosted TTS) and a deterministic multi-source retrieval layer routed across six Model Context Protocol servers, we report three things a research prototype typically lacks. First, a released family of fine-tuned Arabic Islamic model artifacts: an efficient tool-routing LLM (Muslim-6B-PRO, 5.94B parameters) and a Modern Standard Arabic TTS model (Fasih-TTS-V1) that ranks 5th of 17 overall and 2nd of 11 open-weight systems on the community-voted Arabic TTS Arena for MSA. Second, an account and metering layer - a free per-account turn allowance, capacity-aware refusal, and email verification deferred to the point it actually matters - that turns an open demo into an operable, abuse-resistant product. Third, a three-layer observability stack (liveness, error reporting, product analytics) built specifically around the system's characteristic failure mode: a GPU-bound agent host going silent while the web tier keeps serving normally. We report real, measured latency and accuracy figures (98.4% recitation-validation accuracy on 124 cases; end-to-end voice latency of 0.9-1.7s) and discuss the concrete engineering trade-offs and limitations of running an Islamic-knowledge voice product in production.

补充信息

↑