AI 中文总结
本文对比僧伽罗语-泰米尔语电子政务信息检索的CLIR方法,发现多语言嵌入模型BGE-M3性能最优,优于查询翻译方法,可作为低资源政务领域RAG的有效方案。
AI 中文摘要
本文对用于通过僧伽罗语和泰米尔语查询检索英文政府信息的跨语言信息检索(CLIR)方法开展对比评估。研究考察两类CLIR范式:查询翻译(QT),采用Google Translate、NLLB和mBART50;跨语言嵌入(CLE),采用LaBSE、multilingual E5和BGE-M3,同时以单语英文检索作为基线。实验在经人工验证的基准上开展,该基准包含500组僧伽罗语、泰米尔语和英文问答对,源自斯里兰卡政府信息中心(GIC)的1699个分段上下文。采用Recall@k(k=1、3、5、10、15)评估检索性能。单语检索表现较差(Recall@15<10%),而所有CLIR方法均大幅提升检索准确率。其中,BGE-M3取得最高Recall@15,僧伽罗语-英文对达96.2%、泰米尔语-英文对达95.6%,优于最优QT方法(Google Translate对应值为92.4%和93.0%),且可避免翻译开销。这些结果表明,多语言嵌入模型可为低资源政务领域的跨语言检索增强生成(RAG)提供更有效且可扩展的解决方案。
英文摘要
This paper presents a comparative evaluation of cross-lingual information retrieval (CLIR) methods for retrieving English government information using Sinhala and Tamil queries. Two CLIR paradigms are investigated: Query Translation (QT), employing Google Translate, NLLB, and mBART50, and Cross-Lingual Embeddings (CLE), using LaBSE, multilingual E5, and BGE-M3, with monolingual English retrieval as the baseline. Experiments are conducted on a human-verified benchmark comprising 500 Sinhala, Tamil, and English question-answer pairs derived from 1,699 segmented contexts from Sri Lanka's Government Information Center (GIC). Retrieval performance is evaluated using Recall@k (k = 1, 3, 5, 10, 15). Monolingual retrieval performs poorly (Recall@15 <10%), whereas all CLIR approaches substantially improve retrieval accuracy. Among them, BGE-M3 achieves the highest Recall@15, reaching 96.2% for Sinhala-English and 95.6% for Tamil-English, outperforming the best QT approach (Google Translate: 92.4% and 93.0%) while avoiding translation overhead. These results demonstrate that multilingual embedding models provide a more effective and scalable solution for cross-lingual retrieval-augmented generation (RAG) in low-resource government domains.