arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13717cs.CLeess.AS

StreamHear:面向半监督流式语音识别的领域适配伪标签方法

StreamHear: Domain-Adapted Pseudo-Labeling for Semi-Supervised Streaming Speech Recognition

Zefang Liu, Chenyang Zhu, Sangwoo Cho, Xujun Peng, Shi-Xiong Zhang, Sambit Sahu

首次发表
浏览论文内容

中文总结 AI 辅助

针对流式ASR的领域偏移问题,提出StreamHear半监督方法,通过微调离线教师生成伪标签并结合先验正则化重对齐,在四个数据集上性能优于有监督微调且缩小与离线教师的差距。

中文摘要 AI 辅助

流式自动语音识别(ASR)在领域偏移的目标音频上表现不佳,准备标注的领域内数据成本高昂,而未标注音频则十分丰富。本文提出StreamHear,一种半监督流程,通过在标注训练集上微调离线 transducer 教师模型以适配预训练的流式学生模型,在未标注部分生成伪标签,并在混合数据上微调学生模型。我们进一步引入先验正则化的动态规划重对齐步骤,利用ASR假设锚点修正分块级单词位置。在涵盖金融通话、朗读语音和电话质量对话的四个数据集上,StreamHear始终优于有监督学生微调方法,并缩小了与离线教师模型的性能差距。

英文摘要

Streaming automatic speech recognition (ASR) underperforms on domain-shifted target audio, where labeled in-domain data is costly to prepare while unlabeled audio is abundant. We present StreamHear, a semi-supervised pipeline that adapts a pretrained streaming student by fine-tuning an offline transducer teacher on the labeled training set, generating pseudo-labels on the unlabeled portion, and fine-tuning the student on the mixture. We further introduce a prior-regularized dynamic-programming realignment step that fixes chunk-level word placement using an ASR-hypothesis anchor. Across four datasets spanning financial calls, prepared read speech, and phone-quality dialogue, StreamHear consistently outperforms supervised student fine-tuning and narrows the gap to the offline teacher.

发表机构

  • Capital One(第一资本金融公司)

机构由 AI 辅助整理,请以论文原文为准。

↑