发表机构
Recho Inc.(Recho公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出无边界上下文偏置解码器,利用字符级Aho-Corasick自动机与深度门控和阅读空间匹配,在日语和汉语无词边界场景下提升罕见词召回率,并发布首个开放日语基准。
AI 中文摘要
上下文偏置在推理时为ASR系统提供预期词列表,但现有方法依赖日语和汉语不提供的词边界。我们提出一种面向冻结公共CTC模型的无边界偏置解码器,基于字符级Aho-Corasick自动机构建,无需训练且无第二遍解码。两种基于证据的机制替代了词边界:深度自适应门控,根据匹配深度设定推动强度;以及阅读空间匹配,用于音频正确但字符错误的情况。在Aishell-1 NE的困难R1子集上,我们达到66.5%的召回率,高于训练过的CLAS基线(64%),并无需重新调参即可迁移至WenetSpeech和第二种架构。我们发布了首个开放日语上下文偏置基准,在该基准上,偏置在精度超过97%时将罕见词召回率提升25个百分点,面对1000词列表时仍提升19和22个百分点。
英文摘要
Contextual biasing supplies an ASR system with a list of expected words at inference time, but existing methods rely on word boundaries that Japanese and Chinese do not provide. We present a boundary-free biasing decoder for frozen public CTC models, built on a character-level Aho-Corasick automaton, with no training and no second pass. Two evidence-based mechanisms replace the boundary: a depth-adaptive gate that sets how hard to push from match depth, and reading-space matching for when the audio is right but the characters are wrong. On Aishell-1 NE's hard R1 subset we reach 66.5% recall, above the trained CLAS baseline (64%), transferring to WenetSpeech and to a second architecture without retuning. We release the first open Japanese contextual-biasing benchmark, where biasing lifts rare-word recall by 25 points at precision above 97%, and still by 19 and 22 points against 1,000-word lists.