AlgoRAG:面向理论计算机科学教育的检索增强生成——算法分析与复杂性理论的综合评估框架
AlgoRAG: Retrieval-Augmented Generation for Theoretical Computer Science Education -- A Comprehensive Evaluation Framework for Algorithm Analysis and Complexity Theory
浏览论文内容
中文总结 AI 辅助
AlgoRAG通过检索增强生成结合领域知识库,在179道考试题上实现100%成功率,有效提升算法分析与复杂性理论教学。
中文摘要 AI 辅助
教授抽象的计算机科学理论(TCS)概念,如算法分析和复杂性理论,具有挑战性,因为学生必须掌握形式化证明和渐近推理,而传统资源很少以自适应的、按需的方式解释这些内容。我们提出了AlgoRAG,一个专门的检索增强生成(RAG)系统,它将大型语言模型(LLM)与一个精选的、领域特定的知识库相结合,以应对这些挑战。该知识库整合了权威教科书、847张讲座幻灯片、312道带解答的练习题、156个可操作的证明模板和89份复杂性工作表。AlgoRAG融入了领域特定的优化,包括数学实体识别、符号感知检索和教学重排序。我们在179道精选的考试风格问题上评估了AlgoRAG,这些问题涵盖渐近分析、递推关系、动态规划、图算法、NP完全性、排序和分治。该系统实现了100%的成功率,平均响应时间为38.0秒。虽然BLEU-4得分为零——这是n-gram匹配在数学证明上的一个已知局限性,因为等价推理可能使用完全不同的符号——但AlgoRAG取得了ROUGE-1 F1为0.0963、ROUGE-L F1为0.0683以及教学质量得分为0.7620,表明即使表面措辞与参考答案不同,其回答也结构良好且教学上合理。在NP完全性(ROUGE-1 F1 = 0.1285,教学质量 = 0.7643)和图算法(ROUGE-1 F1 = 0.1023,教学质量 = 0.8086)上,性能尤为突出。这些结果支持了RAG是用于个性化理论计算机科学教学的有效架构这一结论,即使对于高度抽象的主题,也能提供正确且上下文丰富的解释。
英文摘要
Teaching abstract theoretical computer science (TCS) concepts such as algorithm analysis and complexity theory is challenging because students must handle formal proofs and asymptotic reasoning that conventional resources rarely explain in an adaptive, on-demand way. We present AlgoRAG, a specialized Retrieval-Augmented Generation (RAG) system that couples a large language model (LLM) with a curated, domain-specific knowledge base to address these challenges. The knowledge base integrates authoritative textbooks, 847 lecture slides, 312 practice problems with solutions, 156 worked proof templates, and 89 complexity worksheets. AlgoRAG incorporates domain-specific optimizations including mathematical entity recognition, notation-aware retrieval, and pedagogical re-ranking. We evaluate AlgoRAG on 179 curated exam-style questions spanning asymptotic analysis, recurrence relations, dynamic programming, graph algorithms, NP-completeness, sorting, and divide-and-conquer. The system achieves a 100% success rate with a mean response time of 38.0 seconds. While BLEU-4 scores are zero -- a known limitation of n-gram matching on mathematical proofs where equivalent reasoning may use entirely different notation -- AlgoRAG attains ROUGE-1 F1 of 0.0963, ROUGE-L F1 of 0.0683, and a pedagogical quality score of 0.7620, indicating that responses are well-structured and didactically sound even when surface wording diverges from reference answers. Performance is especially strong on NP-completeness (ROUGE-1 F1 = 0.1285, pedagogical quality = 0.7643) and graph algorithms (ROUGE-1 F1 = 0.1023, pedagogical quality = 0.8086). These results support the conclusion that RAG is an effective architecture for personalized theoretical-CS instruction, providing correct, context-rich explanations even for highly abstract topics.
发表机构
- Kathmandu University(加德满都大学)
机构由 AI 辅助整理,请以论文原文为准。