用于稳健ASR前端的带学习型观测添加的并行时间-频带混合
Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends
浏览论文内容
中文总结 AI 辅助
提出带学习型观测添加的并行时间-频带混合前端,用于稳健ASR,可降低词错误率且参数量与计算量小。
中文摘要 AI 辅助
语音增强常被用作稳健自动语音识别(ASR)的前端,但循环时序及跨频带模块会引入序列依赖,降低并行效率。本文提出一种基于并行时间-频带混合器(PTBM)模块的序列并行分带增强前端,该模块消除了块内循环展开。PTBM在统一并行架构内整合了带内时序混合与逐帧跨频带注意力,可在时间与频率维度实现高效上下文建模。该系统保留掩码加残差重建接口,并引入学习型观测添加(LOA),无需开发集调优即可抑制ASR敏感伪影。在DNS挑战赛与CHiME-4数据集上,结合冻结的Whisper后端开展的实验表明,所提前端相较循环分带基线可持续降低词错误率,且前端网络仅需0.96M参数与0.58GMAC/s的计算量。
英文摘要
Speech enhancement is often used as a front-end for robust ASR, yet recurrent temporal and cross-band modules introduce sequential dependencies that reduce parallel efficiency. In this paper, we present a sequence-parallel band-split enhancement front-end built on a Parallel Time-Band Mixer (PTBM) block that eliminates within-block recurrent unrolling. PTBM integrates intra-band temporal mixing and per-frame cross-band attention within a unified parallel architecture, enabling efficient contextual modeling across both time and frequency dimensions. The system retains the mask-plus-residual reconstruction interface and introduces learned Observation-Adding (LOA) to suppress ASR-sensitive artifacts without development-set tuning. Experiments on DNS Challenge and CHiME-4 with frozen Whisper back-ends show that the proposed front-end consistently reduces word error rate relative to recurrent band-split baselines while requiring only 0.96 M parameters and 0.58 GMAC/s for the front-end network.
发表机构
- Concordia University(康考迪亚大学)
- McGill University(麦吉尔大学)
- Shenzhen University of Advanced Technology(深圳理工大学)
机构由 AI 辅助整理,请以论文原文为准。