arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向日语音乐搜索查询的低延迟拼写纠错

Low-Latency Spell Correction for Japanese Music Search Queries

Anshul Garg, Pavni Tandon, Karan Bhukar, Tanmay Khandelwal, Ujjal Kumar Dutta

arXiv 2609.04262首次发表:更新:

发表机构

Amazon(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对日语音乐搜索查询的拼写纠错挑战,提出基于BART的紧凑序列到序列模型,结合感知脚本的合成错误生成流程,在低延迟下实现优于基线的拼写纠错性能。

AI 中文摘要

日语搜索查询的拼写纠错面临独特挑战,因为四种书写脚本(拉丁/罗马字、平假名、片假名、汉字)并存,且每种脚本会引发不同的错误模式。我们提出一种紧凑的基于BART的序列到序列模型(3个编码器层+3个解码器层),用于日语音乐搜索查询的低延迟拼写纠错。核心贡献在于一种感知脚本的合成拼写错误生成流程,该流程结合键盘布局模型(QWERTY与 flick输入)、从真实查询日志挖掘的语音混淆先验、清浊辅音交替以及假名大小写错误,生成逼真的训练数据。一个关键设计决策是在合成拼写错误前将混合脚本的目录标题归一化为单一规范脚本,我们证明这对减少模型幻觉至关重要。我们在目标音乐目录上训练自定义的字节级BPE分词器,以在统一词汇表中处理所有四种脚本。在精心整理的评估集上的实验表明,我们的模型达到41.09%的精确匹配准确率和11.62%的字符错误率(CER),优于编辑距离基线,且在单GPU上实现低于4ms的推理延迟,同时在所有评估系统中获得最低字符错误率。我们还通过系统的消融研究分析了各脚本及混合脚本查询的性能,证明了感知脚本的数据增强的有效性。

英文摘要

Spell correction for Japanese search queries presents unique challenges due to the co-existence of four writing scripts (Latin/romaji, hiragana, katakana, and kanji) and the distinct error patterns each script induces. We present a compact BART-based sequence-to-sequence model (3 encoder + 3 decoder layers) designed for low-latency spell correction of Japanese music search queries. The core contribution lies in a script-aware synthetic misspelling generation pipeline that produces realistic training data by combining keyboard-layout models (QWERTY and flick input), phonetic confusion priors mined from real query logs, voiced/unvoiced consonant alternations, and kana case errors. A key design decision is normalizing mixed-script catalog titles to a single canonical script before misspelling synthesis, which we show is critical for reducing model hallucinations. We train a custom byte-level BPE tokenizer on the target music catalog to handle all four scripts in a unified vocabulary. Experiments on a curated evaluation set show that our model achieves an exact-match accuracy of 41.09% and a character error rate (CER) of 11.62%, outperforming edit-distance baselines and achieving the lowest character error rate among all evaluated systems while maintaining sub-4ms inference latency on a single GPU. We further analyze performance across individual scripts and mixed-script queries, demonstrating the effectiveness of script-aware data augmentation through systematic ablation studies.

Comments9 pages, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑