StreamHear:面向半监督流式语音识别的领域适配伪标签方法
StreamHear: Domain-Adapted Pseudo-Labeling for Semi-Supervised Streaming Speech Recognition
浏览论文内容
中文总结 AI 辅助
针对流式ASR的领域偏移问题,提出StreamHear半监督方法,通过微调离线教师生成伪标签并结合先验正则化重对齐,在四个数据集上性能优于有监督微调且缩小与离线教师的差距。
中文摘要 AI 辅助
流式自动语音识别(ASR)在领域偏移的目标音频上表现不佳,准备标注的领域内数据成本高昂,而未标注音频则十分丰富。本文提出StreamHear,一种半监督流程,通过在标注训练集上微调离线 transducer 教师模型以适配预训练的流式学生模型,在未标注部分生成伪标签,并在混合数据上微调学生模型。我们进一步引入先验正则化的动态规划重对齐步骤,利用ASR假设锚点修正分块级单词位置。在涵盖金融通话、朗读语音和电话质量对话的四个数据集上,StreamHear始终优于有监督学生微调方法,并缩小了与离线教师模型的性能差距。
英文摘要
Streaming automatic speech recognition (ASR) underperforms on domain-shifted target audio, where labeled in-domain data is costly to prepare while unlabeled audio is abundant. We present StreamHear, a semi-supervised pipeline that adapts a pretrained streaming student by fine-tuning an offline transducer teacher on the labeled training set, generating pseudo-labels on the unlabeled portion, and fine-tuning the student on the mixture. We further introduce a prior-regularized dynamic-programming realignment step that fixes chunk-level word placement using an ASR-hypothesis anchor. Across four datasets spanning financial calls, prepared read speech, and phone-quality dialogue, StreamHear consistently outperforms supervised student fine-tuning and narrows the gap to the offline teacher.
发表机构
- Capital One(第一资本金融公司)
机构由 AI 辅助整理,请以论文原文为准。