发表机构
Sophea AI Lab; KIEFER SA(索菲亚人工智能实验室; 基弗公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文报告了构建生产级希腊语-英语语音识别系统Sophea的工程实践,通过数据流水线优化和ROVER集成等方法,使系统通过全部九个生产门禁,并显著降低词错误率。
AI 中文摘要
我们报告了一项为期数月的工程项目,旨在构建Sophea,一个生产级双语希腊语-英语自动语音识别系统。我们针对九个生产门禁对该系统进行评估,这些门禁涵盖希腊语和英语的词错误率、语言识别以及非语音音频上的幻觉问题。在二十三次训练迭代和两种模型架构中,没有任何训练数据组合能同时通过全部九个门禁。要满足希腊语嘈杂环境目标,需要约1,500步的密集领域暴露,而保持英语语言识别能力仅容忍约250步,或在使用重新平衡的数据混合(这会降低希腊语准确率)时容忍约1,250步。我们描述了一个六阶段数据流水线,其中针对领域内锚点校准音频质量过滤器,将评分希腊语音频的丢弃比例从98.7%降至10.6%。一项预先注册的消融实验将幻觉缺陷隔离到单个训练数据包中。一个三模型ROVER集成将门禁覆盖率从单个模型的4-7/9提升至9/9,并将重叠语音的WER从53.35%降至37.87%,相对改进29%。一个独立的、针对每个片段学习的仲裁器(基于两个模型)在公开的Open ASR排行榜上列为sophea/asr-k1(预览版),在八个公共英语测试集上平均WER为4.26%,并在实时希腊语嘈杂环境流量上达到25.88%的WER。我们还记录了五个测量工具产生看似合理但错误结果的案例,以及七个已评估但未发布的重大方法。我们不发布任何模型权重或训练数据;我们仅报告方法论和定量结果。
英文摘要
We report a multi-month engineering program to build Sophea, a production bilingual Greek-English automatic speech recognition system. We evaluate the system against nine production gates covering Greek and English word error rate, language identification, and hallucinations on non-speech audio. Across twenty-three training iterations and two model architectures, no training-data composition passed all nine gates simultaneously. Meeting the Greek noisy-environment target required about 1,500 steps of dense domain exposure, while preserving English language identification tolerated only about 250 steps, or about 1,250 with a rebalanced mix that reduced Greek accuracy. We describe a six-stage data pipeline in which calibrating an audio-quality filter against in-domain anchors reduced the discarded share of scored Greek audio from 98.7 percent to 10.6 percent. A pre-registered ablation isolated a hallucination defect to one training-data package. A three-model ROVER ensemble increased gate coverage from 4-7 of 9 for individual models to 9 of 9 and reduced overlapping-speech WER from 53.35 percent to 37.87 percent, a 29 percent relative improvement. A separate learned per-clip arbiter over two models is listed as sophea/asr-k1 (preview) on the public Open ASR Leaderboard, with 4.26 percent average WER across eight public English test sets, and reaches 25.88 percent WER on live Greek noisy-environment traffic. We also document five cases in which a measurement tool produced a plausible but incorrect result and seven substantial approaches that were evaluated but not shipped. No model weights or training data are released; we report methodology and quantitative results only.
Comments21 pages, 6 figures, 4 tables https://huggingface.co/spaces/KIEFERSA/sophea-asr-k1-docs