arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种具有词元级转录歧义容忍度的自动语音识别训练准则

A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition

Saurabh Kumar, Diptiman Mohanta, Prasanta Kumar Ghosh

arXiv 2609.30160首次发表:更新:

发表机构

Indian Institute of Science (IISc)(印度科学研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对转录歧义,提出词元级通配符弧的CTC变体OTC,结合词元与词级逃逸路径及熵索引调度,在19语言25任务上优于CTC,平均相对WER降低9.45%。

AI 中文摘要

自动语音识别通常在假设参考转录是话语唯一有效标注的前提下进行训练,然而即使是名义上逐字转录的文本,也包含发音、拼写或词汇实现上的局部差异,而这些差异在声学上并不能唯一确定。全时态分类(OTC)通过向连接时序分类(CTC)对齐图添加通配符路径来容忍此类噪声,但其词级弧过于粗糙,因为绕过单个不支持的词元会丢弃整个词的监督信息。我们将通配符弧移至词元粒度,使得不支持的词元可以被绕过,而词的其余部分仍保持监督,并将词元级和词级弧结合为互补的逃逸路径。在19种语言和三个语料库上,词元级OTC在所有25项任务上均优于CTC。我们还将通配符权重的基于轮次索引的松弛替换为基于预测熵索引的调度,该调度在性能相当的同时减少了对训练长度的依赖。将此调度与混合图结合,在每个语料库上均获得最低的平均词错误率(WER),相对于CTC实现了9.45%的平均相对WER降低。独立的验证者转录表明,词元级模型在争议字符上放置的通配符绕过概率显著高于CTC,表明词元级容忍度针对的是局部转录歧义。

英文摘要

Automatic speech recognition is typically trained assuming that the reference transcript is the only valid labeling of an utterance, yet even nominally verbatim transcripts contain localized differences in pronunciation, spelling, or lexical realization that the acoustics do not uniquely determine. Omni-temporal Classification (OTC) tolerates such noise by adding wildcard paths to the connectionist temporal classification (CTC) alignment graph, but its word-level arcs are too coarse, since bypassing one unsupported token discards supervision for the whole word. We move wildcard arcs to token granularity so unsupported tokens can be bypassed while the rest of the word stays supervised, and we combine token- and word-level arcs as complementary escape paths. Across 19 languages and three corpora, token-level OTC improves over CTC on all 25 tasks. We also replace epoch-indexed relaxation of the wildcard weights with a predictive-entropy-indexed schedule, which performs comparably while reducing dependence on training length. Combining this schedule with the hybrid graph gives the lowest mean word error rate (WER) on every corpus and a 9.45% average relative WER reduction over CTC. Independent validator transcriptions show that token-level models place significantly more wildcard-bypass probability than CTC on disputed characters, indicating that token-level tolerance targets localized transcript ambiguity.

Comments5 pages, 2 figures, 4 tables; submitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑