探索阿拉伯语中的检索增强生成
Exploring Retrieval Augmented Generation in Arabic
- Newgiza University(新吉萨大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对阿拉伯语RAG应用研究不足的问题,本文通过全面案例研究探索检索阶段的多种语义嵌入模型与生成阶段的若干LLM,还涉及方言差异问题,证实现有模型可有效构建阿拉伯语RAG流水线。
AI中文摘要:
近年来,检索增强生成(Retrieval Augmented Generation, RAG)已成为自然语言处理领域的一项强大技术,它结合了基于检索和基于生成的模型的优势,以增强文本生成任务。然而,RAG在阿拉伯语这一具有独特特征且受资源限制的语言中的应用仍未得到充分探索。本文针对阿拉伯语文本的RAG实现与评估开展了全面的案例研究。该研究重点探索了检索阶段的多种语义嵌入模型,以及生成阶段的若干大语言模型(LLM),以探究在阿拉伯语语境下哪些方案有效、哪些无效。研究还涉及了检索阶段中文档方言与查询方言之间的差异问题。结果表明,现有的语义嵌入模型和LLM可被有效用于构建阿拉伯语RAG流水线。
英文摘要:
Recently, Retrieval Augmented Generation (RAG) has emerged as a powerful technique in natural language processing, combining the strengths of retrieval-based and generation-based models to enhance text generation tasks. However, the application of RAG in Arabic, a language with unique characteristics and resource constraints, remains underexplored. This paper presents a comprehensive case study on the implementation and evaluation of RAG for Arabic text. The work focuses on exploring various semantic embedding models in the retrieval stage and several LLMs in the generation stage, in order to investigate what works and what doesn't in the context of Arabic. The work also touches upon the issue of variations between document dialect and query dialect in the retrieval stage. Results show that existing semantic embedding models and LLMs can be effectively employed to build Arabic RAG pipelines.