arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无边界上下文偏置:面向无分词语言的深度自适应门控与阅读空间匹配

Boundary-Free Contextual Biasing: Depth-Adaptive Gating and Reading-Space Matching for Unsegmented Languages

Muhammad Huzaifah, Yu Pan, Zachary Yeo, Ningjie Bai, Guangzhao Yang

arXiv 2610.09467首次发表:更新:

发表机构

Recho Inc.(Recho公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出无边界上下文偏置解码器,利用字符级Aho-Corasick自动机与深度门控和阅读空间匹配,在日语和汉语无词边界场景下提升罕见词召回率,并发布首个开放日语基准。

AI 中文摘要

上下文偏置在推理时为ASR系统提供预期词列表,但现有方法依赖日语和汉语不提供的词边界。我们提出一种面向冻结公共CTC模型的无边界偏置解码器,基于字符级Aho-Corasick自动机构建,无需训练且无第二遍解码。两种基于证据的机制替代了词边界:深度自适应门控,根据匹配深度设定推动强度;以及阅读空间匹配,用于音频正确但字符错误的情况。在Aishell-1 NE的困难R1子集上,我们达到66.5%的召回率,高于训练过的CLAS基线(64%),并无需重新调参即可迁移至WenetSpeech和第二种架构。我们发布了首个开放日语上下文偏置基准,在该基准上,偏置在精度超过97%时将罕见词召回率提升25个百分点,面对1000词列表时仍提升19和22个百分点。

英文摘要

Contextual biasing supplies an ASR system with a list of expected words at inference time, but existing methods rely on word boundaries that Japanese and Chinese do not provide. We present a boundary-free biasing decoder for frozen public CTC models, built on a character-level Aho-Corasick automaton, with no training and no second pass. Two evidence-based mechanisms replace the boundary: a depth-adaptive gate that sets how hard to push from match depth, and reading-space matching for when the audio is right but the characters are wrong. On Aishell-1 NE's hard R1 subset we reach 66.5% recall, above the trained CLAS baseline (64%), transferring to WenetSpeech and to a second architecture without retuning. We release the first open Japanese contextual-biasing benchmark, where biasing lifts rare-word recall by 25 points at precision above 97%, and still by 19 and 22 points against 1,000-word lists.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑