GUIDE:基于语音与视觉共享ID编码的生成式无监督中文查询纠错方法
GUIDE: Generative Unsupervised Chinese Query Correction via Phonetic and Visual Shared-ID Encoding
浏览论文内容
中文总结 AI 辅助
本文针对中文查询纠错中监督方法维护成本高、无监督方法易意图偏移的问题,提出基于语音与视觉共享ID编码的生成式无监督框架GUIDE,在公开与真实数据集上均优于基线,线上测试也验证了其有效性。
中文摘要 AI 辅助
中文查询纠错(CQC)对内容平台的搜索与查询推荐十分重要,但监督方法依赖大量标注纠错对,随着查询词汇演变,维护成本高昂。无监督纠错借助语言模型颇具吸引力,但在短查询场景下,无约束生成常将歧义输入过度纠正为高频短语,导致意图偏移。本文提出生成式无监督CQC框架\ extsc{GUIDE},采用“混淆-澄清”范式,通过共享ID编码语音或视觉易混淆字符,借助编码器-解码器架构重构原始查询,将纠错约束于合理的混淆邻域,同时从未标注查询流中学习;引入时间衰减、查询频率加权的目标函数,进一步支持对快速变化的查询词汇的适配。在\ extit{QSpell 250K}与大规模真实数据集(\ extit{KwaiSearch})上的实验表明,\ extsc{GUIDE}始终优于强基线,线上A/B测试进一步证实其在纠错质量与下游用户参与度上的提升。
英文摘要
Chinese query correction (CQC) is important for search and query recommendation on content platforms, but supervised methods rely on large annotated correction pairs that are costly to maintain as query vocabularies evolve. Unsupervised correction with language models is attractive, yet in the short-query setting, unconstrained generation often over-corrects ambiguous inputs toward high-frequency phrases, causing intent drift. We propose \textsc{GUIDE}, a generative unsupervised framework for CQC based on a confuse-then-clarify paradigm. \textsc{GUIDE} encodes phonetically or visually confusable characters with shared-IDs and reconstructs the original query with an encoder--decoder architecture, which constrains correction to plausible confusion neighborhoods while learning from unlabeled query streams. A time-decayed, query-frequency-weighted objective further supports adaptation to rapidly changing query vocabularies. Experiments on \textit{QSpell 250K} and a large-scale real-world dataset (\textit{KwaiSearch}) show that \textsc{GUIDE} consistently outperforms strong baselines, while online A/B testing further confirms gains in correction quality and downstream engagement.
发表机构
- Kuaishou Technology(快手科技)
- School of Data Science, Fudan University(复旦大学数据科学学院)
机构由 AI 辅助整理,请以论文原文为准。