arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16379cs.CLcs.SD

未适配的多语种自动语音识别(ASR)在加鲁西库尔德语评估集上:通用参考的分阶段归一化分析

Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis

Hiwa Asadpour

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对未适配的多语种ASR模型在加鲁西库尔德语评估集上的问题,采用通用参考分阶段归一化方法,发现书写系统差异会影响评分,且微调系统表现更差,将发布修正结果供独立核查。

中文摘要 AI 辅助

评估采用拉丁字母领域正字法书写的库尔德语变体的语音识别时,若使用输出阿拉伯字母的模型,会在建模前就产生测量问题:直接评分会将书写系统差异视为识别错误。联合归一化参考文本和假设文本可避免这一问题,但也会改变参考文本的分词方式,使一致性提升与评分分母的变化混杂在一起。本文在未适配的情况下使用发布的MMS-1B-all模型搭配中库尔德语(ckb)适配器,对来自5位说话者的1722段加鲁西问卷语音片段(共9763个参考词标记,时长117.9分钟)进行评估。本文采用通用参考设计:参考文本被折叠一次并固定为9763个标记,仅假设文本表示形式发生变化。原始阿拉伯字母假设文本的词错误率(WER)为111.70%,字符错误率(CER)为100.92%,无精确词匹配;拉丁转写的WER为102.36%,CER为57.89%;将其折叠为参考文本的简化正字法后,WER为97.85%,CER为51.20%。因此,从原始到折叠的转换使测量的WER降低了13.85个百分点,CER降低了49.72个百分点;仅折叠操作就贡献了4.51和6.69个百分点的降幅。仍存在大量错误:14.53%的参考标记为精确匹配,编辑错误以替换为主,且较短片段的WER更高。在相同设计下评估的南库尔德语微调系统(aranemini/southern-kurdish-asr),对每位说话者的表现均更差(1703个片段),其WER为109.56%,CER为55.85%。不过,有12330个输出字符不在折叠表范围内,因此这些错误率必须针对修正后的固定参考文本重新计算。MMS输出还包含613个未转换或未映射的字符,表明部分剩余错误反映了评分流程的限制,而非单纯的识别问题。本文将根据源语料库共享条款发布修正后的参考文本和片段级结果,以支持独立核查。

英文摘要

Evaluating speech recognition for a Kurdish variety written in a Latin field orthography, using a model that outputs Arabic script, creates a measurement problem before a modelling one: direct scoring treats writing-system differences as recognition errors. Jointly normalizing reference and hypothesis avoids this, but also changes reference tokenization, mixing agreement gains with a change in the scoring denominator. I evaluate MMS-1B-all with the Central Kurdish (ckb) adapter, used as released without adaptation, on 1,722 Garrusi questionnaire segments from five speakers (9,763 reference word tokens; 117.9 minutes). I use a common-reference design: the reference is folded once and fixed at 9,763 tokens, while only the hypothesis representation varies. The raw Arabic-script hypothesis scores 111.70% WER and 100.92% CER, with zero exact word matches. Latin transliteration gives 102.36% WER and 57.89% CER; folding it into the reference's reduced orthography gives 97.85% and 51.20%. Thus RAW-to-FOLDED reduces measured WER by 13.85 points and CER by 49.72 points; folding alone accounts for 4.51 and 6.69 points. Substantial error remains: 14.53% of reference tokens are exact matches, edits are substitution-dominated, and per-segment WER is higher for shorter segments. A Southern Kurdish fine-tuned system (aranemini/southern-kurdish-asr), scored under the same design, performs worse on every speaker (1,703 segments), with 109.56% WER and 55.85% CER. However, 12,330 output characters fall outside the folding table, so these rates must be recomputed against the corrected fixed reference. The MMS output also contains 613 unconverted or unmapped characters, showing that part of the residual error reflects scoring-pipeline limits rather than recognition alone. I will release the fixed reference and segment-level results, subject to source-corpus sharing terms, to support independent checking.

补充信息

↑