基于注意力的流式语音识别的选择性前瞻
Selective Lookahead for Attention-Based Streaming ASR
- KU Leuven(荷语鲁汶大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对流式注意力语音识别中前瞻随层数增长的问题,提出有界前瞻块编码器和动态未来块解码,以学习型触发器实现选择性前瞻,在保持准确率的同时大幅降低延迟。
AI中文摘要:
端到端基于注意力的语音识别在离线场景下准确,但难以流式处理:输出可能依赖于未来的音频,且每层少量未来上下文会使前瞻随层数增加而增长。我们通过两种机制解决此问题。一种有界前瞻块编码器通过一个由自注意力和深度卷积共享的年龄选择规则,将每个块的未来感受野限制为恒定数量的块,与层数无关。在该编码器上,动态未来块解码允许逐令牌触发器提交令牌或等待并重新解码;我们提出学习型触发器作为通用机制,并以简单置信度阈值作为有效后备方案。在完整LibriSpeech test-clean上,动态系统以306毫秒的中位延迟匹配最佳静态前瞻准确率(6.5%),而单块静态前瞻为860毫秒,且等待预算将延迟尾部限制在静态基线第90百分位以下,词错误率仅高0.1个百分点。
英文摘要:
End-to-end attention-based speech recognition is accurate offline but hard to stream: outputs can depend on future audio, and a little future context per layer makes the lookahead grow with the number of layers. We address this with two mechanisms. A bounded-lookahead chunk encoder caps every chunk's future receptive field at a constant number of chunks, independent of the number of layers, via one age-selection rule shared by self-attention and the depthwise convolution. On this encoder, dynamic future-chunk decoding lets a per-token trigger commit a token or wait and re-decode it; we propose a learned trigger as the general mechanism, with a simple confidence threshold as an effective fallback. On full LibriSpeech test-clean the dynamic system matches the best static-lookahead accuracy (6.5%) at a median latency of 306 ms versus 860 ms for one-chunk static lookahead, and a wait budget bounds the deferral tail below the static baseline's 90th percentile at 0.1 points more WER.