发表机构
BRAC University(BRAC大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对低资源语言孟加拉语的RAG框架缺陷,提出结构感知RAG框架,通过层级图与图神经网络优化检索,实现更优的MCQ生成与答案预测性能。
AI 中文摘要
传统的检索增强生成(RAG)框架处理文档时未关注其层级结构,导致性能不佳,尤其是在孟加拉语这类低资源语言中。为解决该问题,本文提出一种结构感知RAG框架,将孟加拉语教材建模为层级图,并使用经对比训练的图神经网络检索一小段相关段落。这些段落为大语言模型提供聚焦上下文,支持主题特定的多项选择题(MCQ)生成和领域内答案预测。实验结果表明,该框架在检索指标上优于强大的密集检索基线,生成的MCQ相关性更高,且答案预测准确率更优。
英文摘要
Traditional retrieval-augmented generation (RAG) frameworks process documents without attending to their hierarchical structure, leading to poor performance, especially in low-resource languages such as Bengali. To address this, we propose a structure-aware RAG framework that models Bengali textbooks as hierarchical graphs and uses a contrastively trained graph neural network to retrieve a small set of relevant passages. These passages provide focused context for a large language model, enabling topic-specific multiple-choice question (MCQ) generation and in-domain answer prediction. Experimental results demonstrate that our framework outperforms strong dense retrieval baselines across retrieval metrics, produces more relevant MCQs, and achieves superior answer prediction accuracy.