发表机构
Department of Computer Science
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对孟加拉语KBQA的低资源挑战,提出HybridRAG-BN框架,整合混合检索、答案生成与LoRA微调的验证模块,经实验获竞赛双排行榜第一。
AI 中文摘要
知识库问答(KBQA)系统依赖有效的检索与推理机制,从外部知识源生成准确答案。然而,针对孟加拉语这类低资源语言开发可靠的KBQA系统仍具挑战,原因包括聚焦检索的研究有限、语言资源稀缺,以及生成的响应难以基于外部知识进行落地。本研究提出HybridRAG-BN,这是一个面向孟加拉语KBQA的检索增强框架,整合了基于BM25与BGE-M3的混合检索、使用GGUF格式Gemma-4-31B-Instruct的答案生成,以及经LoRA微调的Gemma-4-31B-Instruct模型用于答案验证与优化。为进一步提升鲁棒性,该框架加入了后处理阶段,通过备选答案替换和DuckDuckGo辅助检索处理未解决的情况。实验结果验证了所提框架的有效性,其在公开排行榜上的令牌级F1分数为0.71654,在私人排行榜上为0.72912,在竞赛中获得第一名。
英文摘要
Knowledge-base question answering (KBQA) systems rely on effective retrieval and reasoning mechanisms to generate accurate answers from external knowledge sources. However, developing reliable KBQA systems for low-resource languages such as Bangla remains challenging due to limited retrieval-focused research, scarce language resources, and difficulties in grounding generated responses in external knowledge. In this work, we propose HybridRAG-BN, a retrieval-augmented framework for Bangla KBQA that integrates hybrid retrieval using BM25 and BGE-M3, answer generation using the GGUF version of Gemma-4-31B-Instruct, and a LoRA-fine-tuned Gemma-4-31B-Instruct model for answer verification and refinement. To further improve robustness, the framework incorporates a post-processing stage that addresses unresolved cases through fallback answer replacement and DuckDuckGo-assisted retrieval. Experimental results demonstrate the effectiveness of the proposed framework, achieving token-level F1 scores of 0.71654 and 0.72912 on the public and private leaderboards, respectively, securing first place in the competition.
CommentsDeveloped for the IEEE Computer Society CUET Student Branch