AI 中文总结
本研究测试了LLMs检索授课视频片段回答编程问题的有效性,对比三种模型与人类讲师,发现专有模型表现接近专家,试点部署获学生良好反馈,为安全整合LLMs到计算机课程提供了可行方案。
AI 中文摘要
本研究评估了利用大语言模型(LLMs)从已录制的授课视频中检索目标片段以回答入门级编程环境中学生问题的有效性。通过限制AI仅识别经教育者验证的现有媒体而非生成开放式文本,该方法旨在缓解生成式幻觉、认知旁路等常见教学风险。我们将三种不同模型与人类讲师的手动视频选择进行基准测试,其中两种为专有模型(Gemini 3.1 Pro和GPT 5.4 Pro),一种为开源权重模型(Qwen 3.5 397B)。随后,一个自动评判框架从相关性、充分性、冗余性及无关材料存在性等维度评估输出结果。尽管AI检索的时间戳与人类基准极少完全重叠,但专有模型在提供充分且高度相关的答案方面达到了与专家近乎相当的水平。此外,在一个规模约为900人的大型C语言课程中对该检索系统进行的试点部署显示出强劲的用户参与度,学生主要利用该工具复习基础概念。通过利用AI检索已确立的授课材料,该方法展现出将LLMs安全整合到新手计算机课程中的可靠、高保真途径的潜力。
英文摘要
This study evaluates the effectiveness of utilising large language models (LLMs) to retrieve targeted segments from delivered video recordings to answer student questions in introductory programming environments. By restricting AI to identifying existing, educator-verified media rather than generating open-ended text, this approach aims to mitigate common pedagogical risks such as generative hallucinations and cognitive bypassing. We benchmarked three distinct models, two proprietary (Gemini 3.1 Pro and GPT 5.4 Pro) and one open-weight (Qwen3.5 397B), against a human lecturer's manual video selections. An automated judging framework subsequently assessed the outputs for relevance, sufficiency, redundancy, and the presence of extraneous material. While the AI-retrieved timestamps rarely shared exact overlaps with the human baseline, the proprietary models achieved near-parity with the expert in delivering sufficient and highly relevant answers. Furthermore, a pilot deployment of this retrieval system in a large C programming cohort (n~=900) demonstrated strong user engagement, with students primarily utilising the tool to review foundational concepts. By leveraging AI to retrieve established lecture material, this approach shows potential for a reliable, high-fidelity pathway for safely integrating LLMs into novice computing courses.
Comments7 pages, 3 tables, 1 figure