arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EduGuard:一种用于编程教育的基于安全检索增强生成的语言模型导师

EduGuard: A Safe RAG-Based LLM Tutor for Programming Education

S M Asif Hossain, Ruksat Khan Shayoni, M. F. Mridha, Jungpil Shin

arXiv 2607.15738首次发表:更新:

AI 中文总结

研究针对学生使用GenAI进行编程学习时的问题,提出EduGuard安全RAG辅导框架,集成多种功能。通过构建基准测试并与强基线比较,在正确性、基础等方面表现最佳,能提升准确率并降低过度依赖,证明安全GenAI辅导需多方面保障。

AI 中文摘要

生成式人工智能(GenAI)越来越多地被学生用于编程解释、调试和作业支持。然而,不受限制的大语言模型(LLM)导师可能会产生幻觉、与课程政策相矛盾、透露完整答案并助长被动依赖。本文提出了EduGuard,一种用于入门编程的安全检索增强生成(RAG)辅导框架。EduGuard集成了查询理解、教师认可的课程检索、教学策略选择、基于评分标准的生成、声明级验证和过度依赖控制。为了使评估来源明确,我们构建了BILearn-CS,这是一个由教师编写、助教验证的包含600个查询的基准,涵盖概念问题、调试案例、误解、作业支持请求、孟加拉语-英语混合查询和对抗性直接答案提示。除了仅基于合成的基准测试,我们还在一个包含150个查询的公开CS50风格课程论坛集上进行评估,并使用平衡的前测/后测设计对10名本科生进行了小型对照试验。使用Meta-Llama-3.1-8B-Instruct作为主要生成器、混合FAISS/BM25检索以及在结构上独立的DeBERTa-v3-large-MNLI作为验证器,将EduGuard与强大的基线进行比较:GPT-4o-mini导师、Llama苏格拉底导师、LPITutor风格的RAG、带有评分标准提示的RAG和同模型自我检查的RAG。在BILearn-CS上,EduGuard在正确性(90.1%)、基础(89.4%)和评分标准对齐(90.8%)方面表现最佳,幻觉率(4.9%)和直接答案泄露率(9.8%)最低。在试验中,相对于GPT-4o-mini导师,它将测试后的即时准确率从68.4%提高到81.2%并将过度依赖从38.0%降低到17.0%。这些结果表明,安全的GenAI辅导不仅需要检索或强大的提示,还需要明确的教学控制、证据验证和部署保障。

英文摘要

Generative AI (GenAI) is increasingly used by students for programming explanation, debugging, and assignment support. Yet unrestricted large language model (LLM) tutors can hallucinate, contradict course policy, reveal complete solutions, and foster passive dependence. This paper presents EduGuard, a safe retrieval-augmented generation (RAG) tutoring framework for introductory programming. EduGuard integrates query understanding, instructor-approved course retrieval, pedagogical strategy selection, rubric-aware generation, claim-level verification, and overreliance control. To make evaluation provenance explicit, we construct BILearn-CS, a 600-query instructor-authored, TA-validated benchmark spanning concept questions, debugging cases, misconceptions, assignment-support requests, code-mixed Bangla-English queries, and adversarial direct-answer prompts. Moving beyond a synthetic-only benchmark, we further evaluate on a 150-query public CS50-style course-forum set and run a small controlled pilot with 10 undergraduates using a counterbalanced pre-test/post-test design. Using Meta-Llama-3.1-8B-Instruct as the primary generator, hybrid FAISS/BM25 retrieval, and DeBERTa-v3-large-MNLI as an architecturally separate verifier, EduGuard is compared against strong baselines: GPT-4o-mini Tutor, Llama Socratic Tutor, LPITutor-style RAG, RAG with rubric prompting, and RAG with same-model self-checking. On BILearn-CS, EduGuard attains the best correctness (90.1%), grounding (89.4%), and rubric alignment (90.8%), with the lowest hallucination (4.9%) and direct-answer leakage (9.8%). In the pilot, it raises immediate post-test accuracy from 68.4% to 81.2% and cuts overreliance from 38.0% to 17.0% relative to GPT-4o-mini Tutor. These results suggest safe GenAI tutoring requires not only retrieval or strong prompting, but explicit pedagogical control, evidence verification, and deployment safeguards.

CommentsAccepted at ACM ICCA 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑