arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HybridRAG-BN:面向孟加拉语知识库问答的带微调验证的检索增强框架

HybridRAG-BN: A Retrieval-Augmented Framework with Fine-Tuned Verification for Bangla KBQA

Rathijit Aich, Nirjhar Das, Mahfuzulhoq Chowdhury

arXiv 2608.13004首次发表:更新:

发表机构

Department of Computer Science

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对孟加拉语KBQA的低资源挑战,提出HybridRAG-BN框架,整合混合检索、答案生成与LoRA微调的验证模块,经实验获竞赛双排行榜第一。

AI 中文摘要

知识库问答(KBQA)系统依赖有效的检索与推理机制,从外部知识源生成准确答案。然而,针对孟加拉语这类低资源语言开发可靠的KBQA系统仍具挑战,原因包括聚焦检索的研究有限、语言资源稀缺,以及生成的响应难以基于外部知识进行落地。本研究提出HybridRAG-BN,这是一个面向孟加拉语KBQA的检索增强框架,整合了基于BM25与BGE-M3的混合检索、使用GGUF格式Gemma-4-31B-Instruct的答案生成,以及经LoRA微调的Gemma-4-31B-Instruct模型用于答案验证与优化。为进一步提升鲁棒性,该框架加入了后处理阶段,通过备选答案替换和DuckDuckGo辅助检索处理未解决的情况。实验结果验证了所提框架的有效性,其在公开排行榜上的令牌级F1分数为0.71654,在私人排行榜上为0.72912,在竞赛中获得第一名。

英文摘要

Knowledge-base question answering (KBQA) systems rely on effective retrieval and reasoning mechanisms to generate accurate answers from external knowledge sources. However, developing reliable KBQA systems for low-resource languages such as Bangla remains challenging due to limited retrieval-focused research, scarce language resources, and difficulties in grounding generated responses in external knowledge. In this work, we propose HybridRAG-BN, a retrieval-augmented framework for Bangla KBQA that integrates hybrid retrieval using BM25 and BGE-M3, answer generation using the GGUF version of Gemma-4-31B-Instruct, and a LoRA-fine-tuned Gemma-4-31B-Instruct model for answer verification and refinement. To further improve robustness, the framework incorporates a post-processing stage that addresses unresolved cases through fallback answer replacement and DuckDuckGo-assisted retrieval. Experimental results demonstrate the effectiveness of the proposed framework, achieving token-level F1 scores of 0.71654 and 0.72912 on the public and private leaderboards, respectively, securing first place in the competition.

CommentsDeveloped for the IEEE Computer Society CUET Student Branch

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑