双形式自动语音识别(DF-ASR):面向中文语音识别的语义感知逆文本规范化
Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition
AI总结:
本文提出DF-ASR框架,通过LLM驱动的双形式监督及ITN-MWER目标,实现语义感知的中文ASR-ITN,在SpeechIO中文子集上性能优于开源系统,可保留转录形式的提示级控制。
AI中文摘要:
现代自动语音识别(ASR)场景既需要忠实转录的口语形式文本,也需要通过逆文本规范化(ITN)得到的可读书面形式文本。然而,这些形式通常由级联模块生成:口语形式ASR的输出会被单独的ITN组件改写,这使得书面形式ASR-ITN易受识别错误影响,且将规范化与声学-上下文建模解耦,尤其针对语义依赖的数值表达式。本文提出双形式自动语音识别(DF-ASR)框架,该框架通过成对的口语形式与书面形式监督,将口语形式ASR能力扩展至语义感知的书面形式ITN,同时保留转录形式的提示级选择能力。双形式监督通过大语言模型(LLM)驱动的“生成-评判”流程构建,训练进一步通过ITN-MWER增强,该序列级目标为规范化敏感跨度的错误分配更高代价。我们还引入决策感知的REQUIRE-ITN/FORBID-ITN协议,分别测量所需规范化与禁止跨度保留情况。在SpeechIO的人工标注中文子集上,DF-ASR始终优于开源ASR-ITN系统,与强大的闭源参考系统性能相当,且能在口语形式与书面形式输出间保持可靠的提示级控制。
英文摘要:
Modern automatic speech recognition (ASR) scenarios require both spoken-form transcripts for faithful transcription and readable written-form transcripts with inverse text normalization (ITN). However, these forms are typically produced by cascaded modules, where a spoken-form ASR output is rewritten by a separate ITN component, making written-form ASR-ITN vulnerable to recognition errors and decoupling normalization from acoustic-contextual modeling, especially for semantically dependent numeric expressions. In this paper, we propose Dual-Form ASR (DF-ASR), a framework that extends spoken-form ASR capability to semantics-aware written-form ITN through paired spoken-form and written-form supervision while retaining prompt-level selection between transcript forms. The dual-form supervision is constructed via a large language model (LLM)-driven generate-and-judge workflow, and training is further enhanced by ITN-MWER, a sequence-level objective that assigns higher cost to errors on normalization-sensitive spans. We also introduce a decision-aware REQUIRE-ITN/\FORBID-ITN protocol to separately measure required normalization and forbidden-span preservation. On manually annotated Chinese subsets from SpeechIO, DF-ASR consistently outperforms open-source ASR-ITN systems, remains competitive with strong closed-source references, and preserves reliable prompt-level control between spoken-form and written-form outputs.