arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13708cs.CLcs.AI

TeachMateGPT:面向科学课程材料的多智能体知识基教学评估生成框架

TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials

  • Ahsanullah University of Science and Technology(阿赫桑乌拉科技大学)
  • American International University - Bangladesh(孟加拉国美国国际大学)
  • Jashore University of Science and Technology(杰索尔科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Fatema Tuj Johora Faria, Mukaffi Bin Moin, M. F. Mridha, Jubayer Al Mahmud

AI总结:

TeachMateGPT作为多智能体知识基框架,通过四项改进解决现有RAG系统局限,生成的科学评估题提升了忠实度与相关性,相关数据集经教师评分验证。

AI中文摘要:

自动生成基于教材的评估题目可减轻科学教师的工作量,但现有检索增强生成(RAG)系统依赖平面检索,仅支持单题生成,缺乏针对弱证据的保障措施,且不适合低资源、板考结构的课程。我们通过TeachMateGPT解决这些局限,该多智能体系统为基于课程的科学评估创作带来四项改进:(i)COPE,一个分层知识库,用多分辨率索引替代令牌窗口分块,沿教学大纲结构分割文档,并通过可遍历基于图的谱系以三种粒度链接它们,使证据与每个主题的教学水平匹配;(ii)分阶段、故障关闭的智能体流水线,替代“先检索后生成”的单次流程:路由门执行搜索,检索在覆盖门的控制下融合密集和词汇证据,在证据不足时暂停生成,专业智能体则起草客观题和构造响应题;(iii)SAVER,一个源归因验证协议,针对检索到的证据对忠实度、相关性和幻觉风险进行评分,对每个创意题的四个子部分应用更严格的接地检查,辅以教师在环评估而非自动过滤;(iv)NCTB-SciGen8,一个基于课程的数据集,包含198道题目(143道选择题、55道创意题),覆盖NCTB八年级科学教材的全部14章,由该流水线生成并经三名在职教师评分。TeachMateGPT相比普通RAG基线,将忠实度从0.68提升至0.96,答案相关性从0.60提升至0.89。

英文摘要:

Automatically generating textbook-grounded assessment items can reduce science teachers' workload, but existing retrieval-augmented generation (RAG) systems rely on flat retrieval, support only single-question generation, lack safeguards against weak evidence, and are ill-suited to low-resource, board-exam-structured curricula. We address these limitations with TeachMateGPT, a multi-agent system contributing four advances to curriculum-grounded science-assessment authoring. (i) COPE, a hierarchical knowledge base replacing token-window chunking with a multi-resolution index that segments documents along syllabus structure and links them at three granularities via a traversable graph-based lineage, matching evidence to each topic's instructional level. (ii) A staged, fail-closed agent pipeline replacing one-shot retrieve-then-generate: routing gates search, retrieval fuses dense and lexical evidence under a coverage gate that withholds generation on insufficient evidence, and specialist agents draft objective and constructed-response items. (iii) SAVER, a source-attributed verification protocol scoring faithfulness, relevance, and hallucination risk against retrieved evidence, applying stricter grounding checks across each creative question's four sub-parts, paired with teacher-in-the-loop evaluation rather than automatic filtering. (iv) NCTB-SciGen8, a curriculum-grounded dataset of 198 items (143 multiple-choice, 55 creative questions) spanning all 14 chapters of the NCTB Class 8 science textbook, produced by the pipeline and rated by three practicing teachers. TeachMateGPT raises faithfulness (0.68 $\rightarrow$ 0.96) and answer relevancy (0.60 $\rightarrow$ 0.89) over a vanilla RAG baseline.

↑