arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14991cs.CLcs.SDeess.AS

Typhoon ASR Streaming:具有实时浅融合的可控低延迟泰语语音识别

Typhoon ASR Streaming: Steerable Low-Latency Thai Speech Recognition with Real-Time Shallow Fusion

  • Typhoon Team, SCB DataX(Typhoon团队,SCB DataX)

机构由 AI 辅助整理,请以论文原文为准。

Warit Sirichotedumrong, Tanawin Samutsin, Shah Faisal Wani, Sittipong Sripaisarnmongkol, Kunat Pipatanakul

AI总结:

提出可部署的流式泰语ASR系统,通过缓存感知编码器和浅融合实现低延迟,在基准上降低字符错误率4.3-4.5倍,解码时控制提升关键词召回率。

AI中文摘要:

开放泰语自动语音识别(ASR)目前由离线、基于Whisper的模型主导,这些模型在转录前需读取整个话语,从而排除了实时字幕和语音代理等低延迟应用。我们提出一个可部署的流式泰语ASR系统,允许用户在解码时控制其词汇表,而无需重新训练。一个广泛使用的、以全上下文训练的开放泰语模型,在作为真正流式运行时性能崩溃;我们通过缓存感知编码器恢复流式能力,通过转换该模型或适配原生流式模型,并添加浅融合层,在流式解码器内部使用GPU n-gram语言模型和短语增强对候选进行重新排序。在两个泰语基准和两种模型规模上,流式模型在全上下文模型失败的情况下保持可用,在一秒前瞻下将字符错误率降低4.3-4.5倍,同时运行速度快于实时。解码时控制将关键词召回率从16.6%提升至20.7%,且无准确率损失,开销可忽略;大部分增益来自对普通训练转录本的n-gram,它解决了模型听到但拼写不一致的代码切换词的书面形式,而短语增强则对罕见领域术语提供有针对性的控制。

英文摘要:

Open Thai automatic speech recognition (ASR) is dominated by offline, Whisper-based models that read the whole utterance before transcribing, ruling out low-latency uses such as live captioning and voice agents. We present a deployable system for streaming Thai ASR that lets a user steer its vocabulary at decode time, without retraining. A widely used open Thai model, trained with full context, collapses when run as a true stream; we restore streaming with a cache-aware encoder, by converting it or adapting a natively streaming one, and add a shallow-fusion layer that re-ranks candidates inside the streaming decoder with a GPU n-gram language model and phrase boosting. Across two Thai benchmarks and two model sizes, the streaming models stay usable where the full-context model fails, cutting character error rate 4.3-4.5x at a one-second look-ahead while running faster than real time. Decode-time steering then lifts keyword recall from 16.6% to 20.7% at no accuracy cost and negligible overhead; most of the gain comes from an n-gram over ordinary training transcripts, which resolves the written form of code-switched words the model hears but spells inconsistently, with phrase boosting adding targeted control over rare domain terms.

补充信息

↑