TurboBias 2.0:面向生产高效型ASR系统的流式上下文偏置技术
TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems
浏览论文内容
中文总结 AI 辅助
本文提出面向生产环境的TurboBias 2.0框架,通过GPU加速的不区分大小写增强图与每流批量解码,在兼顾低延迟高吞吐量的同时,提升了ASR系统的上下文短语识别效果,支持多用户个性化偏置及多种解码方式。
中文摘要 AI 辅助
上下文化对于生产环境下的自动语音识别(ASR)系统至关重要,需在严格延迟约束下准确识别用户提供的短语。尽管诸多上下文偏置方法可提升识别准确率,但往往未满足现代生产ASR系统的实际需求:流式推理、高效批量解码、用户特定上下文列表及低运行时开销。本文提出TurboBias 2.0,这是一种面向生产环境的框架,用于基于Transducer的ASR系统中高效的短语增强。该框架扩展了GPU加速的TurboBias,加入不区分大小写的增强图和每流批量解码,允许批量中的每个 utterance 使用独立的上下文偏置配置,无需共享或混合多用户上下文列表即可实现个性化上下文偏置。该框架支持离线和流式推理,可与贪心和集束搜索解码结合使用。实验表明,TurboBias 2.0在保持低延迟和高吞吐量的同时,提升了上下文短语的识别效果。
英文摘要
Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases must be recognized accurately under strict latency constraints. Although many context-biasing methods improve recognition accuracy, they often do not address the practical requirements of modern production ASR systems: streaming inference, efficient batched decoding, user-specific context lists, and low runtime overhead. We propose TurboBias 2.0, a production-oriented framework for efficient phrase boosting in Transducer-based ASR systems. The framework extends GPU-accelerated TurboBias with a case-insensitive boosting graph and per-stream batched decoding, allowing each utterance in a batch to use an independent context-biasing configuration. This enables personalized context biasing for multiple simultaneous users without sharing or mixing their context lists. The proposed framework supports both offline and streaming inference and can be used with greedy and beam-search decoding. Experiments show that TurboBias 2.0 improves contextual phrase recognition while preserving low latency and high throughput.
发表机构
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。