arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30326cs.SDcs.AI

用于稳健ASR前端的带学习型观测添加的并行时间-频带混合

Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends

Xingyu Shen, Runze Wang, Wei-Ping Zhu, Benoit Champagne

首次发表
浏览论文内容

中文总结 AI 辅助

提出带学习型观测添加的并行时间-频带混合前端,用于稳健ASR,可降低词错误率且参数量与计算量小。

中文摘要 AI 辅助

语音增强常被用作稳健自动语音识别(ASR)的前端,但循环时序及跨频带模块会引入序列依赖,降低并行效率。本文提出一种基于并行时间-频带混合器(PTBM)模块的序列并行分带增强前端,该模块消除了块内循环展开。PTBM在统一并行架构内整合了带内时序混合与逐帧跨频带注意力,可在时间与频率维度实现高效上下文建模。该系统保留掩码加残差重建接口,并引入学习型观测添加(LOA),无需开发集调优即可抑制ASR敏感伪影。在DNS挑战赛与CHiME-4数据集上,结合冻结的Whisper后端开展的实验表明,所提前端相较循环分带基线可持续降低词错误率,且前端网络仅需0.96M参数与0.58GMAC/s的计算量。

英文摘要

Speech enhancement is often used as a front-end for robust ASR, yet recurrent temporal and cross-band modules introduce sequential dependencies that reduce parallel efficiency. In this paper, we present a sequence-parallel band-split enhancement front-end built on a Parallel Time-Band Mixer (PTBM) block that eliminates within-block recurrent unrolling. PTBM integrates intra-band temporal mixing and per-frame cross-band attention within a unified parallel architecture, enabling efficient contextual modeling across both time and frequency dimensions. The system retains the mask-plus-residual reconstruction interface and introduces learned Observation-Adding (LOA) to suppress ASR-sensitive artifacts without development-set tuning. Experiments on DNS Challenge and CHiME-4 with frozen Whisper back-ends show that the proposed front-end consistently reduces word error rate relative to recurrent band-split baselines while requiring only 0.96 M parameters and 0.58 GMAC/s for the front-end network.

发表机构

  • Concordia University(康考迪亚大学)
  • McGill University(麦吉尔大学)
  • Shenzhen University of Advanced Technology(深圳理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑