从多语言流式自动语音识别主干到肯尼亚语言系统:以数据为中心的Nemotron 3.5对基库尤语、多洛语和卡伦金语的适配
From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin
浏览论文内容
中文总结 AI 辅助
该研究针对非洲语言自动语音识别的问题,将NVIDIA Nemotron 3.5 ASR Streaming 0.6B适配到基库尤语、多洛语和卡伦金语,介绍了全参数微调等方法及语料库审核等流程,给出了各语言模型的实验结果及负面发现,为多语言模型适配特定语言系统提供了可审核记录。
中文摘要 AI 辅助
非洲语言的自动语音识别受到拼写不一致、注释工件、音频缺失、说话者和领域不平衡以及与部署不同的评估程序的限制。我们进行了一项端到端工程研究,将NVIDIA Nemotron 3.5 ASR Streaming 0.6B适配到基库尤语、多洛语和卡伦金语。从肯尼亚斯瓦希里语适配的检查点开始,在全参数微调期间保留其缓存感知FastConformer RNN-T、提示条件和流式解码器。该研究涵盖语料库审核、Unicode规范化、拆分检查、持续时间过滤、低速率延续、基于验证的检查点选择、真实流式评估、工件保留和隔离服务。在内部、自适应咨询的评估集上,排除上下文梯度更新,选定的基库尤语和多洛语模型分别实现了42.97%和33.98%的词错误率。多洛语在其冻结历史标签策略下记录了9.59%的字符错误率和8.13%的无空格字符错误率;基库尤语记录了7.79%的无空格字符错误率。卡伦金语仍在进行中:v1-v在一个2411行的干净v3诊断子集上达到68.74%的词错误率,该子集排除了长暂停注释、含数字引用和短于三个令牌的目标。其检查点选择使用了包含测试源行的混合源验证清单,因此该分数不是独立的泛化估计。我们还报告了涉及非语音标签、短话语过度生成、边界敏感词错误率和云作业生命周期故障的负面发现。我们没有声称达到了最新技术水平,因为内部集、重复咨询和规范化与公共基准不同。这项工作提供了一个可审核的记录,说明如何在不放弃流式约束的情况下将多语言流式模型适配到特定语言系统。
英文摘要
Automatic speech recognition (ASR) for African languages is constrained by orthographic inconsistency, annotation artifacts, missing audio, speaker and domain imbalance, and evaluation procedures that differ from deployment. We present an end-to-end engineering study adapting NVIDIA Nemotron 3.5 ASR Streaming 0.6B to Kikuyu, Dholuo, and Kalenjin. Starting from a Kenyan Swahili-adapted checkpoint, we retain its cache-aware FastConformer RNN-T, prompt conditioning, and streaming decoder during full-parameter fine-tuning. The study covers corpus auditing, Unicode normalization, split checks, duration filtering, low-rate continuation, validation-based checkpoint selection, true-streaming evaluation, artifact preservation, and isolated serving. On internal, adaptively consulted evaluation sets excluded from gradient updates at context [56,13], selected Kikuyu and Dholuo models achieve 42.97% and 33.98% WER, respectively. Dholuo records 9.59% CER and 8.13% no-space CER under its frozen historical label policy; Kikuyu records 7.79% no-space CER. Kalenjin remains a work in progress: v1-v reaches 68.74% WER on a 2,411-row clean-v3 diagnostic subset excluding long-pause annotations, digit-bearing references, and targets shorter than three tokens. Its checkpoint selection used a mixed-source validation manifest containing test-origin rows, so the score is not an independent generalization estimate. We also report negative findings involving non-speech labels, short-utterance over-generation, boundary-sensitive WER, and cloud job-lifecycle failures. We make no state-of-the-art claim because the internal sets, repeated consultation, and normalization differ from public benchmarks. This work provides an auditable account of adapting a multilingual streaming model into language-specific systems without discarding streaming constraints.
发表机构
- C-elo Labs(C-elo实验室)
机构由 AI 辅助整理,请以论文原文为准。