arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BengaliMCQ:低资源语言中学术多项选择题的自动生成与答案预测

BengaliMCQ: Automatic Generation and Answer Prediction of Academic Multiple-Choice Questions in a Low-Resource Language

Abu Tarabin Surzo, A. K. M. Nihalul Kabir, Sm Azmain Faysal, Ariana Haque Ami, Lawrence Amlan Gomes, Farig Sadeque

arXiv 2608.15547首次发表:更新:

发表机构

BRAC University(BRAC大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对低资源语言孟加拉语的RAG框架缺陷,提出结构感知RAG框架,通过层级图与图神经网络优化检索,实现更优的MCQ生成与答案预测性能。

AI 中文摘要

传统的检索增强生成(RAG)框架处理文档时未关注其层级结构,导致性能不佳,尤其是在孟加拉语这类低资源语言中。为解决该问题,本文提出一种结构感知RAG框架,将孟加拉语教材建模为层级图,并使用经对比训练的图神经网络检索一小段相关段落。这些段落为大语言模型提供聚焦上下文,支持主题特定的多项选择题(MCQ)生成和领域内答案预测。实验结果表明,该框架在检索指标上优于强大的密集检索基线,生成的MCQ相关性更高,且答案预测准确率更优。

英文摘要

Traditional retrieval-augmented generation (RAG) frameworks process documents without attending to their hierarchical structure, leading to poor performance, especially in low-resource languages such as Bengali. To address this, we propose a structure-aware RAG framework that models Bengali textbooks as hierarchical graphs and uses a contrastively trained graph neural network to retrieve a small set of relevant passages. These passages provide focused context for a large language model, enabling topic-specific multiple-choice question (MCQ) generation and in-domain answer prediction. Experimental results demonstrate that our framework outperforms strong dense retrieval baselines across retrieval metrics, produces more relevant MCQs, and achieves superior answer prediction accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑