arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21084eess.AScs.CL

数字的隐藏代价:ASR系统中的数字归一化与词错误率

The Hidden Cost of Digits: Number Normalization and WER in ASR Systems

  • AGH University of Krakow(克拉科夫AGH大学)

机构由 AI 辅助整理,请以论文原文为准。

Stanisław Kacprzak, Mieszko Fraś

AI总结:

本研究以波兰语为例,分析数字归一化缺失对ASR系统词错误率评估的影响,发现其可造成超过2个百分点的显著差异,甚至超过多语言基准中系统间的差距。

AI中文摘要:

现代自动语音识别(ASR)系统在极大数据集上训练后,可以生成以阿拉伯数字书写的转录文本。这产生了对输出逐字文本的模型进行公平比较以及正确处理参考转录文本的需求。流行的做法通常将文本归一化简化为小写转换和去除标点,对英语以外的语言不进行额外的归一化处理。在本工作中,我们分析了数值表达式归一化对多种语言ASR系统评估的影响,以波兰语这一高度屈折语言为例。我们在VoxPopuli和波兰议会语音数据集上进行了实验,并估算了不同文本归一化方法下的词错误率(WER)差异。我们表明,由于缺乏数字归一化导致的WER差异可能相当大——超过2个百分点,且通常高于流行多语言基准中系统之间的差异。

英文摘要:

Modern automatic speech recognition (ASR) systems trained on extremely large datasets can produce transcripts with numbers written in Arabic numerals. This creates a need for fair comparison with models that output verbatim texts and proper processing of reference transcripts. Popular approaches often reduce text normalization to lowercase and remove punctuation, with no additional normalization applied to languages other than English. In this work, we analyze the impact of normalization of numerical expressions in the evaluation of ASR systems in various languages, using Polish as an example of a highly inflective language. We perform experiments on VoxPopuli and The Polish Parliamentary speech datasets and estimate word error rate (WER) differences for different text normalization approaches. We show that the difference due to the lack of number normalization in WER may be substantial - more than 2 percentage points, and often higher than the differences between systems in popular multilingual benchmarks.

补充信息

↑