KSE-Web:面向低资源高棉语语义搜索的混合检索与大语言模型辅助查询扩展分析
KSE-Web: An Analysis of Hybrid Retrieval and LLM-Assisted Query Expansion for Low-Resource Khmer Semantic Search
浏览论文内容
中文总结 AI 辅助
本文针对低资源高棉语语义搜索,构建专属数据集,评估多种检索方法,发现BM25性能最优,LLM辅助查询扩展未提升效果但模型规模影响显著,为相关研究提供基础。
中文摘要 AI 辅助
高棉语作为低资源语言,存在多项检索挑战,包括标注数据有限、词边界模糊、多语言嵌入模型支持薄弱,以及高棉语与英语频繁混合使用。本文提出KSE-Web,这是一项针对高棉语语义搜索的混合检索与大语言模型(LLM)辅助查询扩展的分析研究。我们从约17000个候选高棉语标题构建数据集,经过滤、归一化、去重和文档长度控制后,保留3000篇清洗后的高棉语文档。该数据集包含300条经人工审核的用户风格高棉语搜索查询,以及经部分人工验证的自动生成相关性标签。我们使用Qwen2.5模型评估字符n元语法BM25、多语言密集检索、BM25+密集混合检索,以及LLM辅助查询扩展。实验结果显示,BM25整体性能最强,召回率达0.943,归一化折损累计增益(nDCG)达0.876;BM25+密集混合检索性能相近,召回率0.929,nDCG 0.871;单独使用密集检索性能较低。LLM辅助查询扩展未优于未扩展检索,但Qwen2.5-3B生成的扩展查询结果远强于Qwen2.5-0.5B,表明LLM规模与扩展质量对低资源高棉语检索至关重要。我们的分析进一步显示,直接LLM扩展会引入主题偏移、通用术语和噪声改写,而简单过滤可能移除有用语义线索。这些发现凸显了LLM辅助高棉语检索的潜力与局限,并为未来构建经更强人工验证标注的高棉语检索数据集及高棉语感知检索模型奠定基础。该数据集及文档将在此HTTP地址提供。
英文摘要
As a low-resource language, Khmer presents several retrieval challenges, including limited annotated data, ambiguous word boundaries, weak support in multilingual embedding models, and frequent mixed Khmer-English usage. This paper presents KSE-Web, an analysis of hybrid retrieval and LLM-assisted query expansion for Khmer semantic search. We construct the dataset from approximately 17K candidate Khmer titles and retain 3K cleaned full-text Khmer documents after filtering, normalization, deduplication, and document-length control. The dataset includes 300 manually reviewed user-style Khmer search queries and silver relevance labels with partial human verification. We evaluate character n-gram BM25, multilingual dense retrieval, hybrid BM25+dense retrieval, and LLM-assisted query expansion using Qwen2.5 models. Experimental results show that BM25 achieves the strongest overall performance, reaching 0.943 Recall and 0.876 nDCG. Hybrid BM25+dense retrieval performs comparably, achieving 0.929 Recall and 0.871 nDCG, while dense retrieval alone performs lower. LLM-assisted query expansion does not outperform non-expanded retrieval; however, Qwen2.5-3B produces substantially stronger expanded-query results than Qwen2.5-0.5B, suggesting that LLM size and expansion quality matter for low-resource Khmer retrieval. Our analysis further shows that direct LLM expansion can introduce topic drift, generic terms, and noisy reformulations, while simple filtering may remove useful semantic cues. These findings highlight both the potential and limitations of LLM-assisted retrieval for Khmer semantic search and provide a foundation for future Khmer retrieval datasets with stronger human-verified annotations and Khmer-aware retrieval models. The dataset and documentation will be made available at github.com/back-kh/KhmerSemantic-Search.