面向德语语音识别的语音学启发式分词:一项跨域研究
Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study
浏览论文内容
中文总结 AI 辅助
本研究对比三类分词器在wav2vec 2.0-CTC模型上的德语跨域语音识别表现,发现分词器性能随词汇量和分布偏移变化,无单一通用最优解。
中文摘要 AI 辅助
德语是一种形态丰富的语言,其音节结构可通过Knuth-Liang断词算法实现极佳预测。本研究探究语音学启发式分词是否可作为端到端语音识别的有竞争力目标。我们在经连接时序分类(CTC)微调的多语种自动语音识别(ASR)wav2vec 2.0主干模型上,对比三类分词器系列:预训练多语种字符集、基于正字法的数据驱动字节对编码(BPE),以及来自Pyphen音节划分和音素转写(grapheme-to-phoneme conversion)的语音学启发式单元。通过40次微调,我们在三类德语测试集上评估,涵盖正交分布偏移:域内朗读语音、方言自发语音、标准德语自发语音。域内场景下,所有语音学启发式分词器在词错误率(WER)和字符错误率(CER)上均与BPE及多语种字符基线相当。在分布偏移下,表现差异沿词汇量而非语言变异轴分化:小词汇量时,感知音节的分词器在方言语音上表现优于BPE,方言语音中语音表层形式多变但音节结构保持,且在自发语音上仍领先,自发语音中新词形式违反词汇闭合规则。音素级混淆分析进一步显示,所有分词器均出现相同的功能词错误,表明声学编码器而非分词器主导错误拓扑。我们的发现表明,分词器选择既取决于部署时预期的分布偏移,也取决于词汇预算,而非存在单一通用最优解。
英文摘要
German is a morphologically rich language whose syllable structure is exceptionally well-predicted by the Knuth--Liang hyphenation algorithm. We ask whether phonologically informed tokenization can serve as a competitive target for end-to-end speech recognition. We compare three tokenizer families on the Omnilingual ASR wav2vec 2.0 backbone fine-tuned with CTC: the pretrained multilingual character inventory, a data-driven Byte-Pair Encoding (BPE) over orthography, and phonologically informed units from Pyphen syllabification and grapheme-to-phoneme conversion. Across 40 fine-tunes, we evaluate on three German test sets spanning orthogonal shifts: in-domain read speech, dialectal spontaneous speech, and standard-German spontaneous speech. In-domain, all phonologically informed tokenizers match BPE and the multilingual character baseline on both WER and CER. Under domain shift the picture splits along vocabulary size rather than the linguistic axis of variation: at small vocabularies, syllable-aware tokenization improves on dialectal speech, where phonetic surface forms vary but syllable structure is preserved, and stays ahead on spontaneous speech, where new word-forms violate vocabulary closure. A phoneme-level confusion analysis further shows that all tokenizers commit the same canonical function-word errors, indicating that the acoustic encoder, not the tokenizer, dominates the error topology. Our findings suggest that tokenizer choice may depend on the vocabulary budget as much as on the distribution shift expected at deployment rather than reducing to a single universal optimum.
发表机构
- Technische Hochschule Nürnberg Georg Simon Ohm(纽伦堡乔治·西蒙·欧姆应用技术大学)
机构由 AI 辅助整理,请以论文原文为准。